Training a decoder on random encoder layers cuts gFID by 27%

single source· 1 articles · confidence: medium · first seen 2026-09-24 20:00 UTC

What this means for you

Nothing to build on today: this is an arXiv preprint, and the abstract mentions no code release. The testable claim is that one decoder handles several layer fusions without retraining; if it holds, changing fusion settings stops requiring a new decoder. Worth reproducing before you plan around it.

FuseReg is a training change for representation autoencoders, which reuse a pretrained visual encoder's features as the compressed latent for both rebuilding and generating images. Which encoder layers form that latent is a trade-off: shallow layers keep pixel detail, deeper ones score better on generation. FuseReg trains over random subsets of layers instead. On ImageNet-256 with DINOv3-L, one decoder handles full, sparse and single-layer fusions without retraining. Swapping in that decoder cut unguided gFID — a generated-image quality score, lower is better — by 27% under an unchanged RAEv2 DiT-XL generator; joint regularisation of both stages cut it 29% on DiT-Base. The scores carry no evaluation date.

Key facts

  • ·FuseReg replaces fixed, heuristic selection of encoder layers with training over random subsets of encoder layers. source
  • ·On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse and single-layer fusions without retraining, achieving higher PSNR than decoders specialised to fixed fusions. source
  • ·Replacing the decoder alone reduced unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. source
  • ·Joint regularisation of both stages reduced unguided gFID by 29% on DiT-Base. source
  • ·Posted as arXiv 2609.31620 on 24 September 2026; the reported scores carry no evaluation date. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire