Command Palette
Search for a command to run...
إلى أي مدى ما زلنا بعيدين عن الاستغناء عن المشفر البصري؟ قوانين التوسع للتدريب المسبق متعدد الوسائط من دون مشفر بصري
إلى أي مدى ما زلنا بعيدين عن الاستغناء عن المشفر البصري؟ قوانين التوسع للتدريب المسبق متعدد الوسائط من دون مشفر بصري
Lin Chen Bolin Ni Qi Yang Lan Jiang Kun Ding Xiaoran Fan Hower Yang Ying Wang Shiming Xiang
الملخص
تعتمد معظم نماذج اللغة الكبيرة متعددة الوسائط (MLLMs) الحديثة على مشفر بصري مدرَّب مسبقًا يوفر سابقًا بصريًا قويًا. في المقابل، تتعلم نماذج MLLMs الخالية من المشفر التمثيلات البصرية مباشرةً من البكسلات الخام، مما يوفر بنية بسيطة وموحدة، لكن سلوكها التوسعي لم يُوصف توصيفًا منهجيًا بعد. ولمعالجة هذه الفجوة، نقارن قوانين التوسع لنماذج MLLMs الخالية من المشفر وتلك المعتمدة على مشفر، ونعرض ثلاث نتائج رئيسية: (1) تؤدي إزالة المشفر البصري إلى إزاحة التخصيص الأمثل للموارد الحاسوبية للهدف متعدد الوسائط نحو نماذج أكبر، بينما تترك التخصيص الخاص بالنص دون تغيير تقريبًا. (2) تُظهر البنيتان حدودًا شبه متداخلة بين الخسارة والحوسبة على الهدف النصي، لكنهما تتباعدان على الهدف متعدد الوسائط: فالنماذج الخالية من المشفر تُظهر أداءً أدنى عند المقاييس الصغيرة، ومن المتوقع أن تلحق بالركب عند نحو 10^22 FLOPs، وهو رقم يقع ضمن ميزانيات التدريب المسبق العملية. (3) في غياب المشفر البصري، يتعلم النموذج اللغوي تولّي دوره عبر تكيف خاص بالرؤية: إذ تصبح التفاعلات ثنائية الاتجاه بين الرموز البصرية متزايدة الفائدة مع نمو الحوسبة التدريبية، وتنتقل المعالجة البصرية نحو الطبقات المبكرة، ويصبح توجيه الخبراء للرموز البصرية أكثر تركيزًا. إجمالًا، تشير نتائجنا إلى أن أفضلية السابق البصري الذي يوفره مشفر مدرَّب مسبقًا تتضاءل مع التوسع، مما يجعل البنى الخالية من المشفر اتجاهًا واعدًا للتدريب المسبق متعدد الوسائط.
One-sentence Summary
Researchers from CASIA, UCAS, and Tencent compare scaling laws for encoder-free and encoder-based multimodal large language models, finding that removing the visual encoder shifts compute-optimal allocation for the multimodal objective toward larger models and that encoder-free models catch up at around 1022 FLOPs through vision-specific adaptation.
Key Contributions
- A systematic scaling-law comparison between encoder-free and encoder-based multimodal LLMs shows that removing the visual encoder shifts compute-optimal decoder allocation toward larger models for the multimodal objective while leaving text-optimal allocation nearly unchanged.
- Encoder-free and encoder-based architectures have nearly overlapping loss-compute frontiers on text, but encoder-free models underperform on the multimodal objective at small scales and are predicted to catch up around 10^22 FLOPs under compute-optimal allocation.
- Encoder-free decoders take over visual encoding as training compute grows, strengthening bidirectional visual-token interactions, shifting visual processing to earlier layers, and concentrating expert routing for visual tokens.
Introduction
Most modern multimodal large language models rely on a pretrained visual encoder to supply the language model with rich visual representations. Encoder-free MLLMs remove this encoder and feed image patches directly into the decoder, which simplifies the architecture but forces the decoder to learn visual representations from raw pixels. Prior encoder-free work has shown feasibility, but its scaling behavior had not been systematically characterized. The authors address this gap with a controlled scaling study comparing encoder-free and encoder-based models that share the same sparse decoder ladder, data mixture, optimization setup, and visual-token granularity. They fit separate scaling laws for text and multimodal objectives and probe decoder internals to show that compute-optimal encoder-free training favors larger models, predict a multimodal efficiency crossover around 10^22 FLOPs, and reveal that the decoder develops vision-specific adaptations such as stronger bidirectional attention among visual tokens and more concentrated expert routing.
Method
The authors establish a framework for estimating scaling laws to determine compute-optimal allocations for large language models. Let M denote the FLOPs per token and D the number of objective tokens, establishing a training budget of C=MD. At a target budget C, the compute-optimal allocation minimizes the loss L(M,D) subject to the constraint MD=C:
Mopt(C),Dopt(C)=M,DargminL(M,D)s.t.MD=C.Across varying budgets, these optima adhere to the compute-optimal allocation law:
Mopt(C)∝Ca,Dopt(C)∝Cb,a+b=1.The corresponding compute-optimal frontiers are defined by:
L∗(C)=E+KC−γ,where γ>0 is the loss-compute exponent, K>0 is a fitted prefactor, and E represents the entropy floor induced by the data distribution. To estimate these parameters, the authors utilize IsoFLOP profiles. For a given budget C, they vary M and set D=C/M, fitting the validation loss as a quadratic function of logM. The fitted minimum identifies Mopt(C), from which Dopt(C) and L∗(C) are directly derived.
In parallel, the authors design a matched model ladder to compare architectural paradigms in multimodal large language models. This ladder comprises 11 sparse mixture-of-experts language models ranging from 1.1B to 44B total parameters and 71M to 2.4B active non-embedding parameters, all sharing the same data mixture, optimization setup, and visual-token granularity. The architectural comparison between encoder-based and encoder-free approaches is illustrated below.
In the encoder-based configuration, images are processed using a pretrained SigLIP 2 Vision Transformer, followed by a ConvPool adapter and a projector. This architecture applies causal attention uniformly to all tokens. The Vision Transformer maintains a consistent size across different decoder scales and is trained jointly with the decoder. Conversely, the encoder-free model bypasses the dedicated visual encoder, mapping raw image patches directly into the decoder through a patch projection. In this setup, visual tokens attend bidirectionally within each image, while all other attention mechanisms remain strictly causal. The authors also investigate a fully causal variant of this encoder-free design to further isolate the impact of bidirectional visual attention.
Experiment
This set of experiments compares encoder-free and encoder-based multimodal models using IsoFLOP profiles and compute-optimal plus overtraining scaling analyses. Removing the pretrained visual encoder leaves text-objective allocation and loss behavior nearly unchanged but shifts multimodal training toward larger decoders and creates an initial multimodal efficiency gap that is projected to close within practical compute budgets, with earlier catch-up on language-heavy topics like STEM and later on perception-heavy topics like captioning. The decoder appears to compensate by developing encoder-like visual processing: visual tokens receive stronger bidirectional attention, are transformed earlier in shallow layers, and are handled by more concentrated expert routing.
Encoder-free models allocate compute to model scale much more aggressively on the multimodal objective than on the text objective. The multimodal allocation exponent remains elevated under causal attention over visual tokens, indicating that the shift toward larger models is not specific to bidirectional visual attention. Encoder-free text allocation exponents stay close to encoder-based text behavior under both attention settings. Multimodal allocation exponents rise substantially relative to text, indicating a compute-optimal preference for larger models. Causal attention over visual tokens yields a multimodal allocation exponent close to the bidirectional setting, so the larger-model shift persists across visual attention masks.
The experiments assess how encoder-free models allocate compute to model scale across text and multimodal objectives under bidirectional and causal visual attention. Multimodal objectives consistently favor substantially larger models than text objectives, while encoder-free text allocation remains close to encoder-based text behavior in both attention settings. The larger-model shift persists under causal visual attention, indicating that this preference is not specific to bidirectional visual attention.