$npx skillfedfor your agent
RESEARCH

Randomly dropping encoder layers during training beats fixed fusion heuristics

on: FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Representation autoencoders face a structural tension: the encoder layers that best preserve pixel detail are not the same ones that make a diffusion model's job easy. RAEv2 documented this as a Pareto frontier — expanding the fused layer subset to include shallower layers raises reconstruction PSNR but also raises gFID. The standard response has been to pick a fixed heuristic and live with the compromise. FuseReg refuses that compromise by treating layer selection as a training distribution rather than a design decision.

The mechanism is straightforward. During training, the decoder and diffusion transformer each receive a randomly sampled, normalized subset of frozen encoder layers instead of a fixed fusion. Because the subset mean is normalized by the number of retained layers, the expected latent equals the full-layer mean regardless of drop rate — no signal is destroyed in expectation. What changes is the variance: sampling adds noise precisely along directions where encoder layers disagree. The paper proves this exactly for linear squared-loss consumers, showing the induced objective contains an explicit cross-layer disagreement penalty. Deterministic global fusion cannot replicate this second-order structure — a learned gate in the experiments collapses toward the shallowest, most pixel-aligned layer once adversarial training begins, which is exactly the shortcut FuseReg is designed to prevent.

The empirical payoff is concrete. A single FuseReg decoder trained on DINOv3-L handles full, sparse, and single-layer fusions without retraining, matching or exceeding specialized decoders on their own training fusions. Swapping only the decoder while keeping the generator and sampled latents fixed reduces unguided gFID on the native fusion, with larger gains when the decoder renders latents from a shifted fusion space. Joint regularization of both decoder and diffusion transformer on DiT-Base yields improvements the paper describes as complementary — the combined gain at the selected configuration exceeds what either stage achieves alone, though the authors note this is not a formal statistical interaction test.

Scaling to DiT-XL introduces nuance. The generator's preferred drop rate becomes metric- and guidance-dependent: the configuration minimizing gFID differs from the one maximizing Inception Score, and internal guidance creates an interaction between prediction heads that a single regularization strength cannot simultaneously optimize. The paper recommends choosing decoder and generator rates separately and is explicit that the linear theory does not predict the sign of two-rate interactions in nonlinear models.

Generalization experiments with SigLIP2-L and EUPE-B confirm that reconstruction robustness and generation gains are not artifacts of DINOv3-L's feature geometry. The approach requires no architectural changes and no additional training budget. The main open questions are how rate preferences shift at higher resolutions and over longer training schedules — both explicitly listed as untested.

Randomizing which encoder layers form the latent during training turns a fixed architectural choice into a regularizer that benefits both reconstruction and generation without extra cost.

Sources & links