$npx skillfedfor your agent
RESEARCH

Training a language model on protein geometry boosts general reasoning, but unevenly

on: Does Learning Protein Folding Generalize to Broader Reasoning?

Post-training a language model on protein structure questions improves its performance on benchmarks that contain no proteins, no structural inputs, and no specialized modules. That is the central claim, and the paper backs it with enough instrumentation to take seriously.

The mechanism is two-part. FoldingCorpus applies twelve deterministic geometric operators to known protein coordinates, generating verified question-answer pairs about contacts, distances, orientations, and chirality. Fold2Reason then post-trains Qwen3.5-9B on those discrete answers through the model's native language head while simultaneously routing the same adapted representations through a frozen coordinate decoder that supplies continuous geometry supervision. At evaluation time, the decoder and all protein-specific machinery are discarded; only the LoRA adapter remains.

The aggregate result is a macro-average gain from 45.09% to 48.33% across ten general-purpose benchmarks spanning spatial, graph, chemistry, and scientific reasoning. All ten datasets show positive mean changes across three training seeds. The ablations are the most informative part: the discrete FoldingCorpus answers account for most of the transfer, while the geometry signal adds a smaller increment concentrated in 3D-oriented tasks. Matched controls trained on random geometry targets, copied answer codes, or shuffled labels produce substantially smaller or negative gains, which rules out the simplest confounds.

The paper is unusually candid about what remains unresolved. Two spatial benchmarks, FTB-Core and SpatialViz, together account for a disproportionate share of the gain, and within those, single tasks like sparse-constraint candidate selection and CubeAssembly dominate. The authors audit answer-selection bias explicitly: the base model selects option D on 79 of 80 CubeAssembly questions despite no reference answer being D. Removing both spatial benchmarks entirely still leaves a positive eight-dataset gain, but the honest reading is that the improvement is task-dependent, not uniform.

The protein-count scaling curve peaks around 2,000 proteins and then shows diminishing returns, while increasing label density from 3 to 12 questions per protein at fixed coverage produces no monotonic improvement. The paper treats this as evidence that data breadth matters more than label density, though it cannot fully separate breadth from compute exposure.

Gemma-4-12B-IT shows no gain, while InternVL3.5-8B and all three Qwen3.5 scales do. Architecture dependence is real and unexplained. Whether the transfer reflects genuinely reusable spatial computation or more reliable invocation of capabilities already latent in the base model is, the authors acknowledge, unresolved. The protein folding quality itself remains far below competitive structure prediction; this is a reasoning-transfer study, not a folding system.

Protein structure supervision transfers measurable reasoning gains to a general LM, but the effect is architecture-dependent and task-concentrated enough to demand careful reading of the ablations.

Sources & links