TL;DR: A representation autoencoder builds its latent by averaging layers of a frozen vision encoder, and which layers to use is usually picked by hand. FuseReg instead trains the decoder and the diffusion model on a random mix of layers at every step. The resulting decoder reconstructs well from many different mixes, and generated images get better, with no change to the model or its inference cost.
A fixed layer choice forces a trade-off. Shallow encoder layers keep pixel detail, which helps reconstruction; deeper layers are easier for the diffusion model to generate. With RAEv2 on DINOv3-L, using all 23 layers instead of the last 7 raises reconstruction PSNR from 22.58 to 27.04 dB, but worsens unguided gFID from 1.65 to 3.01.
Training on random layer mixes gives one decoder for many mixes. The same FuseReg decoder reconstructs from all 23 layers, from the last 7, or from a single layer, with higher PSNR than decoders trained for any one of them.
Generation improves too. Keeping a trained diffusion model and its samples fixed and only swapping in the FuseReg decoder lowers unguided gFID from 3.01 to 2.21. Training the diffusion model the same way brings DiT-Base from 13.96 to 9.93, and the gains carry over to the SigLIP2 and EUPE encoders.
A representation autoencoder (RAE[1]) generates images in the feature space of a pretrained vision encoder such as DINOv3[3]. The encoder stays frozen; a diffusion model learns to generate encoder features, and a decoder turns those features back into pixels. RAEv2[2] builds the latent by fusing several encoder layers, and picks the layers by hand.
That choice matters because the decoder and the diffusion model want different things. Shallow layers keep fine pixel detail, which helps the decoder reconstruct the image. Deeper layers are more semantic and tend to be easier for the diffusion model to generate. One fixed fusion makes both stages live with the same compromise.
FuseReg removes the fixed choice during training. Each training example uses the average of a randomly chosen subset of layers, so the decoder learns to reconstruct from many different mixes instead of one. The diffusion model is trained the same way: it sees a random mix and learns to predict the all-layer average. At test time both use the all-layer average, so the architecture, sampler and inference cost stay the same.
For every training image, each of the K encoder layers is dropped at random with probability p, at least one layer is kept, and the kept layers are averaged. Every layer is equally likely to be kept, so on average the result equals the all-layer average; what changes from sample to sample is only where the layers disagree with each other. The model therefore learns not to depend on the quirks of any one layer. The latent also always includes a fixed summary of the final layer (its token mean, marked L23 surrogate in the figures).
The decoder and the diffusion model each get their own drop rate, pdec and pdit. A rate of 0 means no regularization, which is the RAEv2 baseline.
Unless noted, experiments use ImageNet at 256×256 and a frozen DINOv3-L encoder with 23 layers. Reconstruction is measured on 50,000 images with PSNR and SSIM (higher is better) and rFID (lower is better). Generation is measured on 50,000 samples, drawn with 50 Euler steps, with gFID (lower is better) and Inception Score, IS (higher is better).
Each RAEv2 decoder is trained on one fixed fusion: the last 7 layers (K = 7) or all 23 (K = 23). We test every decoder on both fusions and on layer 11 alone.
To isolate the decoder, we take a trained RAEv2 diffusion model, generate one fixed set of 50,000 latents, and decode the same latents with decoders trained at increasing pdec. Any change in image quality comes from the decoder alone. We do this for an RAEv2 generator trained on all 23 layers and one trained on the last 7; all decoders are trained on the 23-layer pool, so the plain decoder is mismatched with the 7-layer generator.
Next we also train the diffusion model on random layer mixes. The grids cover every combination of decoder rate (columns) and generator rate (rows). The green box is the RAEv2 baseline; blue cells beat it and orange cells are worse.
We repeat the DiT-Base experiment with two other frozen encoders: SigLIP2-L[4], a vision-language encoder with 23 layers, and EUPE-B[5], a smaller ViT-B encoder with 11 layers. The training recipe is the same, with rates 0, 0.3, 0.6 and 0.9. With both encoders, FuseReg decoders also reconstruct much better from partial inputs than the fixed-fusion decoder (paper, Appendix E).
Why does this work? The encoder never changes, so the improvement comes from how the decoder reads it. To see this, we measure how much the decoder depends on each layer. FuseReg spreads what it needs across the whole encoder instead of leaning on a few shallow layers.
Which encoder layers a representation autoencoder uses does not have to be decided once and fixed. Training on random layer mixes gives a decoder that works with many mixes and a diffusion model that generates better images. This holds for DINOv3-L at two model sizes and for the SigLIP2-L and EUPE-B encoders, with the encoder, architecture and inference cost unchanged. All experiments are on ImageNet at 256×256.
@misc{du2026fuseregregularizinglayerfusion,
title={FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders},
author={Hongyang Du and Yunfei Xie and Junjie Ye and Jiawei Yang and Xiaoyan Cong and Haodong Zhang and Yongchao Huang and Haiyu Wu and Zongxia Li and Shihang Gui and Dawei Liu and Runhao Li and Jingcheng Ni and Chen Wei and Randall Balestriero and Yue Wang},
year={2026},
eprint={2609.31620},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.31620},
}