Future-frame supervision that vanishes at inference is InternW0-Δ's smartest design choice
The central bet in InternW0-Δ is that future video frames should shape a policy's representations during training without ever appearing at inference. That sounds like a minor implementation detail, but it resolves a genuine tension: video pretraining gives a model rich knowledge of how scenes evolve, yet actually generating future frames at test time is too slow and too expensive for real-time robot control. The paper's answer is Causal Imprint—a set of learnable tokens inside the video expert that are supervised by differences between consecutive future video latents during training, then used to condition the action expert at inference without any future rollout. The ablation numbers make the mechanism's value concrete: adding the direct latent-difference objective moves LIBERO-Plus success from 69% to roughly 71%, and adding the complementary future-feature alignment objective pushes it further to about 76%.
The architecture pairs a pretrained video expert (initialized from Wan2.2-TI2V-5B) with a randomly initialized action expert through 30 Mixture-of-Transformers blocks. A frozen vision-language model supplies scene semantics to the action expert—a deliberate split, because T5 embeddings preserve the video expert's pretrained language-video interface while the VLM handles the observational grounding T5 cannot provide. Swapping VLM backbones in the ablation produces an 8.7-point swing in success rate, which is a surprisingly large sensitivity to a frozen component.
A third training-only signal, 4D-aware distillation from a frozen Track4World teacher, injects geometric and motion priors into the video expert via an auxiliary MSE loss on cached descriptors. The teacher and its student branch are discarded at inference entirely, adding no latency. The gain is modest—roughly 2 percentage points on LIBERO-Plus—but the design principle is clean: pay the compute cost once offline, cache the descriptors, and let the loss shape the video expert's hidden states without touching the inference graph.
The data side is where the scale claim lives. Over 20K hours of processed training data spans 15 robot demonstration datasets, egocentric human video converted through an Ego2Robot pipeline, and UMI data—all mapped into an 80-dimensional canonical action space. The data ablations are instructive: raw egocentric video adds almost nothing to gripper-based benchmarks, while the Ego2Robot conversion and UMI data provide meaningful gains. The paper is candid that this may reflect the current benchmark's gripper focus rather than a fundamental limit of human demonstration data.
On the engineering side, the inference stack is worth noting. Starting from a 780 ms round-trip on the standard runtime, a sequence of optimizations—process isolation, feature caching, context caching, grouped action execution, CUDA Graph replay—brings latency down to 152.8 ms on a single RTX 5090, enabling 30 Hz asynchronous control. The paper reports that LIBERO-Plus success rate drops only marginally across these optimizations, from 92.78% to 92.23%, suggesting the speed gains are not purchased with accuracy.
The honest limitations section flags two gaps: the contribution of egocentric data is not systematically studied across scales or dexterous tasks, and the GPT-assisted correction experiments are preliminary. Both are real gaps, not false modesty.
A disciplined architecture that turns future-video supervision into inference-free predictive representations, backed by the largest open robot training corpus assembled to date.
Sources & links
Related on SkillFed
A corpus-scale audit of 40,285 publicly listed agent skills finds bursty publication, a category-level supply-demand mismatch, heavy intent redundancy, and real risk from…
A survey of the agent skills ecosystem finds 26.1% of community-contributed skills carry a vulnerability, script-bundling doubles the odds, and one operator accounts for over half…