HyperAIHyperAI

Command Palette

Search for a command to run...

視覚エンコーダの除去まであとどのくらいか:エンコーダフリー・マルチモーダル事前学習のスケーリング則

Lin Chen Bolin Ni Qi Yang Lan Jiang Kun Ding Xiaoran Fan Hower Yang Ying Wang Shiming Xiang

概要

現代のほとんどのマルチモーダル大規模言語モデル(MLLM)は、強力な視覚事前分布をもたらす事前学習済み視覚エンコーダの上に構築されている。これに対しエンコーダフリーMLLMは、生のピクセルから直接視覚表現を学習し、単純で統一的なアーキテクチャを提供するが、そのスケーリング挙動は体系的に特徴づけられていない。この空白を埋めるため、本研究ではエンコーダフリーMLLMとエンコーダベースMLLMのスケーリング則を比較し、以下の3つの主要な知見を報告する。(1)視覚エンコーダを除去すると、マルチモーダル目的に対する計算最適な配分はより大規模なモデル側へ移動する一方、テキスト目的に対する配分はほぼ変化しない。(2)2つのアーキテクチャはテキスト目的に関する損失–計算量フロンティアがほぼ重なるが、マルチモーダル目的では乖離する。エンコーダフリーモデルは小規模では劣るものの、約10^22 FLOPsで追いつくと予測され、これは実用的な事前学習予算の範囲内に十分収まる。(3)視覚エンコーダがない場合、言語モデルは視覚特有の適応を通じてその役割を肩代わりすることを学習する。すなわち、学習計算量が増加するにつれて視覚トークン間の双方向相互作用の利益が大きくなり、視覚処理はより初期の層へ移行し、視覚トークンに対する専門家ルーティングはより集中するようになる。総じて、本研究の結果は、事前学習済みエンコーダが与える視覚事前分布の利点が規模とともに減少することを示しており、エンコーダフリーアーキテクチャはマルチモーダル事前学習の有望な方向性として位置づけられる。

One-sentence Summary

Researchers from CASIA, UCAS, and Tencent compare scaling laws for encoder-free and encoder-based multimodal large language models, finding that removing the visual encoder shifts compute-optimal allocation for the multimodal objective toward larger models and that encoder-free models catch up at around 102210^{22}1022 FLOPs through vision-specific adaptation.

Key Contributions

  • A systematic scaling-law comparison between encoder-free and encoder-based multimodal LLMs shows that removing the visual encoder shifts compute-optimal decoder allocation toward larger models for the multimodal objective while leaving text-optimal allocation nearly unchanged.
  • Encoder-free and encoder-based architectures have nearly overlapping loss-compute frontiers on text, but encoder-free models underperform on the multimodal objective at small scales and are predicted to catch up around 10^22 FLOPs under compute-optimal allocation.
  • Encoder-free decoders take over visual encoding as training compute grows, strengthening bidirectional visual-token interactions, shifting visual processing to earlier layers, and concentrating expert routing for visual tokens.

Introduction

Most modern multimodal large language models rely on a pretrained visual encoder to supply the language model with rich visual representations. Encoder-free MLLMs remove this encoder and feed image patches directly into the decoder, which simplifies the architecture but forces the decoder to learn visual representations from raw pixels. Prior encoder-free work has shown feasibility, but its scaling behavior had not been systematically characterized. The authors address this gap with a controlled scaling study comparing encoder-free and encoder-based models that share the same sparse decoder ladder, data mixture, optimization setup, and visual-token granularity. They fit separate scaling laws for text and multimodal objectives and probe decoder internals to show that compute-optimal encoder-free training favors larger models, predict a multimodal efficiency crossover around 10^22 FLOPs, and reveal that the decoder develops vision-specific adaptations such as stronger bidirectional attention among visual tokens and more concentrated expert routing.

Method

The authors establish a framework for estimating scaling laws to determine compute-optimal allocations for large language models. Let MMM denote the FLOPs per token and DDD the number of objective tokens, establishing a training budget of C=MDC = MDC=MD. At a target budget CCC, the compute-optimal allocation minimizes the loss L(M,D)\mathcal{L}(M, D)L(M,D) subject to the constraint MD=CMD = CMD=C:

Mopt(C),Dopt(C)=arg⁡min⁡M,DL(M,D)s.t.MD=C.M_{\mathrm{opt}}(C), D_{\mathrm{opt}}(C) = \underset{M, D}{\arg \min} \mathcal{L}(M, D) \quad \text{s.t.} \quad MD = C.Mopt​(C),Dopt​(C)=M,Dargmin​L(M,D)s.t.MD=C.

Across varying budgets, these optima adhere to the compute-optimal allocation law:

Mopt(C)∝Ca,Dopt(C)∝Cb,a+b=1.M_{\mathrm{opt}}(C) \propto C^a, \qquad D_{\mathrm{opt}}(C) \propto C^b, \qquad a + b = 1.Mopt​(C)∝Ca,Dopt​(C)∝Cb,a+b=1.

The corresponding compute-optimal frontiers are defined by:

L∗(C)=E+KC−γ,\mathcal{L}^*(C) = E + KC^{-\gamma},L∗(C)=E+KC−γ,

where γ>0\gamma > 0γ>0 is the loss-compute exponent, K>0K > 0K>0 is a fitted prefactor, and EEE represents the entropy floor induced by the data distribution. To estimate these parameters, the authors utilize IsoFLOP profiles. For a given budget CCC, they vary MMM and set D=C/MD = C/MD=C/M, fitting the validation loss as a quadratic function of log⁡M\log MlogM. The fitted minimum identifies Mopt(C)M_{\mathrm{opt}}(C)Mopt​(C), from which Dopt(C)D_{\mathrm{opt}}(C)Dopt​(C) and L∗(C)\mathcal{L}^*(C)L∗(C) are directly derived.

In parallel, the authors design a matched model ladder to compare architectural paradigms in multimodal large language models. This ladder comprises 11 sparse mixture-of-experts language models ranging from 1.1B to 44B total parameters and 71M to 2.4B active non-embedding parameters, all sharing the same data mixture, optimization setup, and visual-token granularity. The architectural comparison between encoder-based and encoder-free approaches is illustrated below.

In the encoder-based configuration, images are processed using a pretrained SigLIP 2 Vision Transformer, followed by a ConvPool adapter and a projector. This architecture applies causal attention uniformly to all tokens. The Vision Transformer maintains a consistent size across different decoder scales and is trained jointly with the decoder. Conversely, the encoder-free model bypasses the dedicated visual encoder, mapping raw image patches directly into the decoder through a patch projection. In this setup, visual tokens attend bidirectionally within each image, while all other attention mechanisms remain strictly causal. The authors also investigate a fully causal variant of this encoder-free design to further isolate the impact of bidirectional visual attention.

Experiment

This set of experiments compares encoder-free and encoder-based multimodal models using IsoFLOP profiles and compute-optimal plus overtraining scaling analyses. Removing the pretrained visual encoder leaves text-objective allocation and loss behavior nearly unchanged but shifts multimodal training toward larger decoders and creates an initial multimodal efficiency gap that is projected to close within practical compute budgets, with earlier catch-up on language-heavy topics like STEM and later on perception-heavy topics like captioning. The decoder appears to compensate by developing encoder-like visual processing: visual tokens receive stronger bidirectional attention, are transformed earlier in shallow layers, and are handled by more concentrated expert routing.

Encoder-free models allocate compute to model scale much more aggressively on the multimodal objective than on the text objective. The multimodal allocation exponent remains elevated under causal attention over visual tokens, indicating that the shift toward larger models is not specific to bidirectional visual attention. Encoder-free text allocation exponents stay close to encoder-based text behavior under both attention settings. Multimodal allocation exponents rise substantially relative to text, indicating a compute-optimal preference for larger models. Causal attention over visual tokens yields a multimodal allocation exponent close to the bidirectional setting, so the larger-model shift persists across visual attention masks.

The experiments assess how encoder-free models allocate compute to model scale across text and multimodal objectives under bidirectional and causal visual attention. Multimodal objectives consistently favor substantially larger models than text objectives, while encoder-free text allocation remains close to encoder-based text behavior in both attention settings. The larger-model shift persists under causal visual attention, indicating that this preference is not specific to bidirectional visual attention.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています