Encoder-free multimodal models are on track to match ViT-based ones at practical scale
The pretrained visual encoder's advantage over raw-pixel decoders shrinks predictably with compute, and this paper quantifies exactly when it disappears.
The study runs a controlled ladder of eleven sparse MoE models ranging from 1.1B to 44B total parameters, comparing encoder-based models (using a fixed SigLIP 2 ViT throughout) against encoder-free models that project raw image patches directly into the decoder. Both families share the same data mixture, optimizer, and visual-token count, so the comparison isolates the representation attached to each token rather than anything else.
Three findings stand out. First, removing the encoder shifts compute-optimal allocation toward larger models on the multimodal objective: the model allocation exponent rises meaningfully, while the text exponent barely moves. Encoder-free training simply needs more decoder capacity to handle visual representation learning alongside language modeling. Bootstrap intervals confirm the difference is positive across all replicates.
Second, the multimodal loss frontiers diverge over the measured range but converge in slope. Encoder-free loss falls faster with compute, and extrapolating the fitted scaling laws places the crossover around 10²³ FLOPs under compute-optimal allocation—roughly three orders of magnitude below the pretraining compute of recent flagship models such as Kimi K2.5. Under overtraining the crossover arrives later but remains within the same order of magnitude. Every bootstrap replicate produces a finite crossover.
Third, the decoder doesn't passively absorb visual tokens—it reorganizes to encode them. Bidirectional attention among visual tokens becomes increasingly valuable at scale. Visual token representations diverge from their layer-zero inputs much earlier in encoder-free models than in encoder-based ones, while text token trajectories stay nearly identical across both architectures. Expert routing for visual tokens also grows more concentrated, consistent with a subset of MoE experts taking over the vision-specific role the ViT previously held. The paper calls this vision-specific adaptation, and the probes are specific enough to be convincing.
The crossover timing varies by topic. STEM questions, which lean heavily on language and symbolic reasoning, approach parity fastest. GUI, OCR, and captioning—tasks demanding fine spatial or textual perception—lag considerably, which makes intuitive sense: those are exactly the domains where a pretrained encoder's prior is hardest to replicate from scratch.
The honest caveat is that the crossover prediction extrapolates well beyond the measured compute range and holds the visual encoder fixed at roughly 400M parameters. Jointly scaling the encoder would change the allocation problem entirely. Still, the mechanistic probes give the extrapolation more credibility than a curve-fit alone would warrant.
Encoder-free multimodal models are predicted to match encoder-based ones around 10²³ FLOPs—and the decoder's internal reorganization explains why.
Sources & links
Related on SkillFed
Across 15 LLMs and 1,141 real skills, routing accuracy decays logarithmically as libraries grow. The same slope also predicts execution-side rescue, and editing the library cuts…
A training curriculum that progressively withdraws skill files teaches Qwen2.5-VL agents to internalize procedural knowledge — the resulting policy beats a skill-augmented RL…