Fuse builds social-advice tests where the right answer is scripted

single source· 1 articles · confidence: medium · first seen 2026-09-14 20:00 UTC

What this means for you

Fuse and its 21,000-example dataset are public, so teams building advice-giving assistants can test them for the failure the paper describes: drifting toward however the user framed the situation. No hosted service and no leaderboard — you run the evaluation yourself.

Fuse, a multi-agent simulation (several language models in interacting roles), is released with a 21,000-example dataset to test social advice from assistants. A target agent holds a hidden motive; other agents interact with it, one standing in for the user, who then asks the assistant under test to infer it. Because the authors script that motive, the correct answer is known by construction. A human study with 24,000 annotations checked the simulations were faithful. Testing 12 models, the authors report that user retelling makes social inference harder, that models shift with biased framing and often need more detail than humans do, and that longer conversations do not always help.

Key facts

  • ·Fuse is a multi-agent simulation in which a target agent's hidden motive is set by the researchers, so a correct answer exists by construction. source
  • ·Simulation faithfulness was validated in a human study with 24,000 annotations. source
  • ·The framework was applied to 12 LLMs. source
  • ·Fuse and a dataset of 21,000 examples are open-sourced. source
  • ·Reported findings include sensitivity to biased user framing, a need for more detail than humans require, and no consistent gain from longer conversations. source
  • ·The paper was posted on 14 September 2026 as arXiv 2609.17496. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire