$npx skillfedfor your agent
RESEARCH

Toy block puzzles train better spatial reasoning than real-world annotations

on: SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

Spatial reasoning in vision-language models fails at a surprisingly low level. When shown a simple block structure and asked to predict its 2D projection, leading proprietary models produce wrong answers that humans solve with 95% accuracy. SpatialBlock-15k starts from that gap and asks whether the right training signal is not more real-world annotation but something simpler: the kind of structured block play that developmental psychology links to early spatial cognition.

The dataset contains 15,000 synthetic problems across three task types. The first asks models to project a 3D block arrangement into a 2D view, requiring occlusion reasoning. The second tests whether a model can track a specific block through rotations and flips. The third presents two separate structures and asks what emerges when they are joined at a marked contact point. Each type also has a color-cued variant where hue encodes depth ordering, anchor identity, or overlap location — not decoration, but functional geometry.

The color extension turns out to matter substantially. Removing it drops performance by roughly 8 percentage points on MindCube and around 5 points on MMSI-Bench for the reasoning-trained variant. The mechanism is interpretable: color gives the model a reference object to anchor multi-image reasoning, mimicking how humans parse complex scenes by fixing on a salient element.

Two training strategies are tested. Direct prediction fine-tunes on answer supervision alone. The reasoning variant uses LoRA initialization followed by GRPO reinforcement learning with a three-part reward covering answer correctness, chain-of-thought format, and response length. Full-parameter cold-start fine-tuning was tried and abandoned: even GPT-5 generates unreliable block descriptions, making teacher-distilled trajectories noisy enough to hurt rather than help.

The generalization results are the real argument. Trained entirely on synthetic blocks, the 7B model improves over its backbone by 17.6% on MindCube and outperforms SpaceR and Spatial-SSRL on MMSI-Bench using only 15k samples. The 4B model reaches 51.3% on MindCube, the best reported open-source figure. Performance on MMMU, a general visual benchmark, stays stable, so the spatial gains do not come at the cost of broader perception.

A controlled comparison reinforces the point. A dataset of the same size built around conventional spatial question types — relative direction and distance — adapted into the same block environment improves only on in-domain tasks and degrades elsewhere. The block-stacking formulation, not synthetic data per se, drives the transfer.

The viewpoint robustness check is worth noting: training uses a fixed rendering angle, yet accuracy holds when the same structures are rendered from two other angles. The model appears to have learned something about 3D structure rather than memorizing a particular projection.

Synthetic block puzzles teach spatial reasoning that transfers to real scenes better than real-scene annotation does.

Sources & links

SkillFed lets your AI agent find skills for you

example · real query, live index
agent > wish: “Large Vision-Language Models”
No install? Search from any chat →