Replacing a generator's internal code with geometry features halves camera-trajectory error
single source· 1 articles · confidence: medium · first seen 2026-09-20 20:00 UTC
What this means for you
Nothing to act on: no code, no weights, no API. If you work on scene-consistent video generation, the comparison is worth reading because only the latent changed — generator and training protocol held fixed — so it isolates the representation from the model.
A preprint reparameterises the features of a geometry foundation model into a compact latent space — the numeric code a generator conditions on — instead of adding geometry as another output. That latent decodes jointly to appearance, depth, camera pose and point maps. With the generator and training protocol held fixed, swapping it in cuts FVD, a video-fidelity distance, by 12.7% on RealEstate10K and 23.1% on DL3DV, and halves camera-trajectory error on RealEstate10K. The work, posted to arXiv on 20 September 2026, mentions no code or weights.
Key facts
- ·The paper is arXiv preprint 2609.24981, posted 20 September 2026. source
- ·The geometry-native autoencoder (GAE) latent decodes jointly to appearance, depth, cameras and point maps. source
- ·With the generator and training protocol held fixed, FVD falls 12.7% on RealEstate10K and 23.1% on DL3DV. source
- ·Camera-trajectory error is halved on RealEstate10K. source
- ·No code or weights release is mentioned in the abstract. source
What the sources say
- Hugging Face Daily Papers (research) — Abstract-only report of a representation swap; no code, weights or outside replication indicated.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersGAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation2026-09-20