HyperAIHyperAI

Command Palette

Search for a command to run...

How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

Lin Chen Bolin Ni Qi Yang Lan Jiang Kun Ding Xiaoran Fan Hower Yang Ying Wang Shiming Xiang

Abstract

Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss–compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around 10^22 FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.

One-sentence Summary

Researchers from CASIA, UCAS, and Tencent compare scaling laws for encoder-free and encoder-based multimodal large language models, finding that removing the visual encoder shifts compute-optimal allocation for the multimodal objective toward larger models and that encoder-free models catch up at around 102210^{22}1022 FLOPs through vision-specific adaptation.

Key Contributions

  • A systematic scaling-law comparison between encoder-free and encoder-based multimodal LLMs shows that removing the visual encoder shifts compute-optimal decoder allocation toward larger models for the multimodal objective while leaving text-optimal allocation nearly unchanged.
  • Encoder-free and encoder-based architectures have nearly overlapping loss-compute frontiers on text, but encoder-free models underperform on the multimodal objective at small scales and are predicted to catch up around 10^22 FLOPs under compute-optimal allocation.
  • Encoder-free decoders take over visual encoding as training compute grows, strengthening bidirectional visual-token interactions, shifting visual processing to earlier layers, and concentrating expert routing for visual tokens.

Introduction

Most modern multimodal large language models rely on a pretrained visual encoder to supply the language model with rich visual representations. Encoder-free MLLMs remove this encoder and feed image patches directly into the decoder, which simplifies the architecture but forces the decoder to learn visual representations from raw pixels. Prior encoder-free work has shown feasibility, but its scaling behavior had not been systematically characterized. The authors address this gap with a controlled scaling study comparing encoder-free and encoder-based models that share the same sparse decoder ladder, data mixture, optimization setup, and visual-token granularity. They fit separate scaling laws for text and multimodal objectives and probe decoder internals to show that compute-optimal encoder-free training favors larger models, predict a multimodal efficiency crossover around 10^22 FLOPs, and reveal that the decoder develops vision-specific adaptations such as stronger bidirectional attention among visual tokens and more concentrated expert routing.

Method

The authors establish a framework for estimating scaling laws to determine compute-optimal allocations for large language models. Let MMM denote the FLOPs per token and DDD the number of objective tokens, establishing a training budget of C=MDC = MDC=MD. At a target budget CCC, the compute-optimal allocation minimizes the loss L(M,D)\mathcal{L}(M, D)L(M,D) subject to the constraint MD=CMD = CMD=C:

Mopt(C),Dopt(C)=arg⁡min⁡M,DL(M,D)s.t.MD=C.M_{\mathrm{opt}}(C), D_{\mathrm{opt}}(C) = \underset{M, D}{\arg \min} \mathcal{L}(M, D) \quad \text{s.t.} \quad MD = C.Mopt​(C),Dopt​(C)=M,Dargmin​L(M,D)s.t.MD=C.

Across varying budgets, these optima adhere to the compute-optimal allocation law:

Mopt(C)∝Ca,Dopt(C)∝Cb,a+b=1.M_{\mathrm{opt}}(C) \propto C^a, \qquad D_{\mathrm{opt}}(C) \propto C^b, \qquad a + b = 1.Mopt​(C)∝Ca,Dopt​(C)∝Cb,a+b=1.

The corresponding compute-optimal frontiers are defined by:

L∗(C)=E+KC−γ,\mathcal{L}^*(C) = E + KC^{-\gamma},L∗(C)=E+KC−γ,

where γ>0\gamma > 0γ>0 is the loss-compute exponent, K>0K > 0K>0 is a fitted prefactor, and EEE represents the entropy floor induced by the data distribution. To estimate these parameters, the authors utilize IsoFLOP profiles. For a given budget CCC, they vary MMM and set D=C/MD = C/MD=C/M, fitting the validation loss as a quadratic function of log⁡M\log MlogM. The fitted minimum identifies Mopt(C)M_{\mathrm{opt}}(C)Mopt​(C), from which Dopt(C)D_{\mathrm{opt}}(C)Dopt​(C) and L∗(C)\mathcal{L}^*(C)L∗(C) are directly derived.

In parallel, the authors design a matched model ladder to compare architectural paradigms in multimodal large language models. This ladder comprises 11 sparse mixture-of-experts language models ranging from 1.1B to 44B total parameters and 71M to 2.4B active non-embedding parameters, all sharing the same data mixture, optimization setup, and visual-token granularity. The architectural comparison between encoder-based and encoder-free approaches is illustrated below.

In the encoder-based configuration, images are processed using a pretrained SigLIP 2 Vision Transformer, followed by a ConvPool adapter and a projector. This architecture applies causal attention uniformly to all tokens. The Vision Transformer maintains a consistent size across different decoder scales and is trained jointly with the decoder. Conversely, the encoder-free model bypasses the dedicated visual encoder, mapping raw image patches directly into the decoder through a patch projection. In this setup, visual tokens attend bidirectionally within each image, while all other attention mechanisms remain strictly causal. The authors also investigate a fully causal variant of this encoder-free design to further isolate the impact of bidirectional visual attention.

Experiment

This set of experiments compares encoder-free and encoder-based multimodal models using IsoFLOP profiles and compute-optimal plus overtraining scaling analyses. Removing the pretrained visual encoder leaves text-objective allocation and loss behavior nearly unchanged but shifts multimodal training toward larger decoders and creates an initial multimodal efficiency gap that is projected to close within practical compute budgets, with earlier catch-up on language-heavy topics like STEM and later on perception-heavy topics like captioning. The decoder appears to compensate by developing encoder-like visual processing: visual tokens receive stronger bidirectional attention, are transformed earlier in shallow layers, and are handled by more concentrated expert routing.

Encoder-free models allocate compute to model scale much more aggressively on the multimodal objective than on the text objective. The multimodal allocation exponent remains elevated under causal attention over visual tokens, indicating that the shift toward larger models is not specific to bidirectional visual attention. Encoder-free text allocation exponents stay close to encoder-based text behavior under both attention settings. Multimodal allocation exponents rise substantially relative to text, indicating a compute-optimal preference for larger models. Causal attention over visual tokens yields a multimodal allocation exponent close to the bidirectional setting, so the larger-model shift persists across visual attention masks.

The experiments assess how encoder-free models allocate compute to model scale across text and multimodal objectives under bidirectional and causal visual attention. Multimodal objectives consistently favor substantially larger models than text objectives, while encoder-free text allocation remains close to encoder-based text behavior in both attention settings. The larger-model shift persists under causal visual attention, indicating that this preference is not specific to bidirectional visual attention.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp