Benchmark and model test sound-guided navigation for embodied agents
single source· 1 articles · confidence: medium · first seen 2026-09-19 20:00 UTC
What this means for you
Nothing to act on. No weights, no code and no release date are given, and the benchmark's availability is not stated, so none of the reported numbers can be checked yet. Watch the navigation split: the paper claims parity with vision-only navigation there without printing comparison scores.
A preprint posted on 19 September 2026 introduces OmniEchoBench, a test suite for spatial audio-visual perception and sound-guided navigation, plus OmniEcho, a model built for both. The benchmark spans six tasks, 197 real-world scenes, 2,972 question-answer pairs and 900 navigation samples recorded as first-order ambisonics — audio that carries the direction a sound arrived from, not just its loudness — across 30 environments. The authors report that their model leads on spatial audio-visual perception, and that its sound-guided navigation approaches conventional vision-language navigation. Fine-grained localisation and distance estimation stay unsolved. The abstract mentions no released weights or code.
Key facts
- ·OmniEchoBench contains 197 real-world spatial audio-visual scenes and 2,972 question-answer pairs. source
- ·It includes 900 navigation samples recorded as first-order ambisonics from 30 real-world environments. source
- ·The benchmark defines six tasks covering spatial audio-visual perception and audio-vision-language navigation. source
- ·The authors' OmniEcho model combines a first-order ambisonics spatial encoder with a pretrained semantic audio pathway. source
- ·The paper reports state-of-the-art performance on spatial audio-visual perception without giving comparison scores in the abstract. source
- ·The work was posted to arXiv as a preprint on 19 September 2026. source
What the sources say
- Hugging Face Daily Papers (research) — Abstract only: a new benchmark and model, with no external evaluation and no release details.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersOmniEcho: Spatial Audio Understanding for Embodied Agents2026-09-19