← arXiv

arXiv·3 min read

FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

Choosing encoder layers for shared latent space in representation autoencoders forces a trade‑off between detail preservation and generative performance.

Representation autoencoders reuse features from a pretrained visual encoder as both reconstruction and diffusion latents, but they must decide which encoder layers constitute the shared latent space for the pixel decoder and the generator. Shallow layers retain fine pixel details, whereas deeper layers improve generation metrics, creating a tension that prior work resolves with a fixed heuristic fusion.

FuseReg replaces the heuristic selection by training the downstream decoder on randomly sampled subsets of encoder layers, effectively regularizing against cross‑layer disagreement. This subset sampling forces the model to be insensitive to which layers are fused, allowing a single decoder to reconstruct from full, sparse, or single‑layer fusions without retraining, and extending to diffusion training with joint regularization of both stages.

On ImageNet‑256 using DINOv3‑L, a FuseReg‑trained decoder achieves higher PSNR than decoders specialized to a fixed fusion, and swapping in this decoder reduces unguided gFID by 27% while keeping the RAEv2 DiT‑XL generator unchanged. Applying the same regularization jointly during diffusion training further cuts unguided gFID by 29% on a DiT‑Base model.

TakeawayTraining downstream models for layer‑fusion robustness narrows the reconstruction‑generation gap without modifying the pretrained encoder.

Prodigy briefing — continue on the original for source material, discussion, and updates.

Read the paper ↗