$npx skillfedfor your agent
RESEARCH

Robots gain nothing from longer memory unless their video backbone is autoregressive

on: Long-WAM: Scaling the Context of World-Action Models

The core claim here is deceptively simple: having access to visual history and knowing how to use it are not the same thing. Long-WAM makes this concrete by showing that extending a robot policy's context window from zero to 19.2 seconds raises success on RoboCasa GR-1 from 63.3% to 78.7% — but only when the video backbone was pretrained autoregressively. A bidirectionally pretrained initialization given the same longer context shows no net gain, hovering around 61–64% regardless of window length. The AR advantage actually widens with context: at 19.2 seconds, the robot-domain AR variant leads the bidirectional one by 17.1 points, versus just 3.3 points without any history.

The mechanism is staged. First, a video model called LongLive2.0-Robot trains on roughly 10,000 window-equivalent hours of robot and egocentric footage using autoregressive prediction — no action labels, no shared action space required across embodiments. Then, during world-action adaptation, that causal temporal structure is preserved rather than discarded. Actions are conditioned on both observed history and partially denoised future latents, with the video expert predicting forward before the action expert denoises. This predict-then-act path (called IDM) outperforms joint co-denoising on LIBERO-Long: 99.5% versus 97.8%, with the gap largest on the longest tasks.

Longer context is only useful if the robot can still respond in time. The paper addresses this directly with a deployment stack that combines asynchronous execution, streaming VAE encoding, NVFP4 quantization, and device-specific kernel tuning. On an RTX 5090, the full pipeline — including future-video latent prediction — runs in 107.4 ms per action chunk. That is faster than Fast-WAM at 244.1 ms, despite Long-WAM doing strictly more computation. Crucially, asynchronous switching barely degrades performance: IDM drops only 0.2 percentage points from synchronous to asynchronous on RoboTwin 2.0, while LingBot-VA loses 45.5 points under the same condition.

The real-robot results are the sharpest demonstration. On a Unitree G1, Long-WAM grasps cups from a conveyor moving at 7.5 cm/s in 90% of trials and stacks a moving cup into another in 95% of trials. Both baseline methods — Fast-WAM and a prior policy — succeed zero times on the stacking task across 20 trials each. On a YAM bimanual manipulator, tasks averaging over 40 seconds complete at 81.7% success across three categories.

One honest caveat the paper raises: the decline at 38.4 seconds of context (success falls back to around 75%) coincides with 80.4% of sampled frames being padding, because current robot datasets rarely contain trajectories that long. The memory limit may be a data coverage problem, not an architectural one. Each context length also requires its own trained model variant — a single policy that adapts its window at test time remains future work.

Autoregressive video pretraining is what makes longer robot memory pay off — access to history without that prior is nearly worthless.

Sources & links