Mira-Scene reports 39.8% better object placement than SAM3D

single source· 1 articles · confidence: medium · first seen 2026-09-19 20:00 UTC

What this means for you

Nothing to build on yet. No code, weights or API is stated, and the reported figures carry no evaluation date, so they are not comparable with numbers you already have. If you work on single-image scene reconstruction, the transferable idea is the bounded, pixel-aligned prediction target — but you cannot test it from this paper alone.

Mira-Scene, a method posted to arXiv in September 2026, rebuilds a 3D scene from a single photograph by predicting where each visible object pixel sits in that object's own coordinate space, then aligning those points to a point cloud. That replaces the usual approach: guessing each object's position and orientation as free numbers, which generalises poorly. The authors report a 39.8% relative gain in 3D-IoU (volumetric overlap) and 16.5% in 2D-IoU over the SAM3D baseline, on indoor, outdoor, synthetic and in-the-wild scenes, trained on limited open-source data. It is a preprint, with no evaluation date, harness or code release stated.

Key facts

  • ·Mira-Scene reports a 39.8% relative improvement in 3D-IoU over the SAM3D baseline. source
  • ·It reports a 16.5% relative improvement in 2D-IoU over the same baseline. source
  • ·Evaluation covers indoor, outdoor, synthetic and in-the-wild scenes, trained on limited open-source data. source
  • ·Layout is learned without scene-level layout annotations, using object-level 3D data via a bounded Canonical Coordinate Map. source
  • ·Object transformations are recovered by geometric alignment between the Canonical Coordinate Map and a monocular scene-space Point Cloud Map. source
  • ·The paper is arXiv preprint 2609.23796, posted 19 September 2026; no evaluation date, harness or code release is stated. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire