$npx skillfedfor your agent
RESEARCH

FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation

The core problem FrameMorrow solves is a subtle mismatch in how long-horizon video generators handle memory. Most existing approaches ask: which past frames look most like what's on screen right now? That question is wrong. What matters is which past frames will be useful for what comes next — and those two sets of frames often differ substantially.

FrameMorrow's answer is prospective tokens: a small set of learned representations, predicted autoregressively from available history, recent context, and the known rollout condition, that stand in for future information needs without actually generating future video. Four tokens, predicted causally, capture most of the benefit. These tokens then score eligible historical frames, and the top-ranked frames get passed to the frozen generator through whatever reference-conditioning interface it already exposes.

The training signal is clean: a frozen DINOv2 encoder compares candidate historical frames against realized future frames from training trajectories, producing ranking targets the prospective tokens learn to approximate. At inference, no future frames are available — the tokens do the work alone. A diagnostic in the appendix confirms the teacher's rankings correlate meaningfully with actual downstream generation utility, not just visual similarity.

Because FrameMorrow outputs explicit frames rather than model-internal states, it works across generators it was never trained with, including closed-source ones. On Seedance 2.0 and Kling O3, accessed only through public reference-conditioning APIs, it improves consistency by roughly 1.3 points on both models compared to uniform historical sampling under the same reference budget.

The breadth of evaluation is notable: five benchmarks, eleven generative models, covering long-video generation, interactive single- and multi-shot generation, and action-conditioned world models. On 60-second MovieGenBench generation, FrameMorrow achieves the best average rank across all three tested backbones. With Self-Forcing, imaging quality rises from 62.22 to 68.47 and dynamic degree from 51.72 to 64.56. In a blinded user study across 900 judgments per criterion, raters preferred FrameMorrow for consistency 62.7% of the time versus 16.7% for native models.

The ablations are instructive. Selecting two frames with FrameMorrow outperforms uniform sampling with sixteen frames on consistency — relevance matters more than quantity. Future-grounded supervision adds 3.10 consistency points over a no-future control that uses only recent-context similarity for ranking targets. The gap to an oracle with direct future access narrows from 4.40 to 1.30 points.

Training costs are modest: 12 GPU-hours per selector on two A100s, using 10K videos. Inference overhead is small enough to report in a table without embarrassment. The selector shares one checkpoint across compatible backbones within each condition modality, so the same trained module serves multiple generators.

Selecting historical frames by future relevance rather than present similarity is the right framing, and the results across eleven models make a credible case for it.

Sources & links