skillfed
RESEARCH

Video models commit to wrong answers in early denoising steps and never self-correct

on: VGI-Bench: Probing Visual Intelligence in Video Generation Models

The best video generation model tested here, Seedance 2.0, scores just over half the maximum possible under VGI-Bench's criteria. That number is the headline, but the more interesting finding is what it reveals about why current models fail at procedural visual tasks.

VGI-Bench is a benchmark built around a specific claim: that existing evaluations of video model reasoning are compromised by three problems. First, they use line-art or schematic inputs that sit outside the natural-image distribution these models were trained on, so failures may reflect domain mismatch rather than reasoning limits. Second, most visual reasoning tasks can be answered by inspecting the final frame alone, without requiring the model to actually simulate a valid intermediate trajectory. Third, many benchmark tasks are either trivially easy or hopelessly beyond current capability, making failure uninformative.

The benchmark addresses these by using photorealistic inputs, filtering tasks to those where intermediate state evolution genuinely matters, and running a pre-generation sanity check that accepts only tasks at least one current model can partially solve and at least one fails. Tasks span four domains — Visual Organization, Physical Manipulation, Structured Puzzles, and Spatiotemporal Dynamics — each annotated with skill tags like Topology, Planning, and Affordance.

The evaluation design is notably careful. Two metrics are multiplied together: a global Completeness score (did the video make meaningful progress toward the goal?) and a Rubric Score (did it respect intermediate process constraints?). The multiplicative combination means a video that reaches a plausible final state by cheating the procedure still scores poorly. An adaptive coarse-to-fine frame sampling strategy catches transient rule violations that uniform sampling would miss.

The diagnostic analyses are where the paper earns its keep. Oracle prompting — giving models the explicit step-by-step solution in the prompt — improves performance only modestly, even for the strongest closed-source systems. Visual style matters substantially: switching from realistic to line-art inputs reshuffles model rankings, with open-source models especially sensitive. Synthetic fine-tuning on a million abstract examples transfers to realistic tasks, but only where the training distribution structurally overlaps with the test tasks; gains on non-overlapping tasks are small and sometimes negative.

The denoising trajectory analysis is the most technically pointed section. Decoding intermediate states from four open-source models, the authors find that self-correction — revising a wrong intermediate state into a correct one — occurs at or below a few percent in every step pair and essentially stops after the earliest denoising steps. Wrong-to-wrong transitions are an order of magnitude more common. The model's answer is largely committed early; later steps refine rather than reconsider. This directly challenges claims that iterative denoising in video models resembles the search-like reasoning observed in discrete diffusion language models.

The benchmark's stated scope is narrow by design: image-to-video only, fixed 16:9 aspect ratio, English prompts, and task durations calibrated to the five-to-ten second generation window of current models. Long-horizon procedural reasoning is explicitly out of scope. These are honest constraints, not oversights.

The strongest model tested barely clears half the maximum score, and the denoising analysis shows why: video models commit to wrong answers early and refine rather than correct them.

Sources & links

Live matches from SkillFed’s research index — a weak match is labeled, never suppressed, so an empty-looking result never falsely means “no such research exists.”

SkillFed lets your AI agent find skills for you

example · real query, live index
agent > wish: “video generation models”
No install? Search from any chat →