Rule-induction tests find reasoning models fail on equivalent task variants

single source· 1 articles · confidence: medium · first seen 2026-09-11 20:00 UTC

What this means for you

Nothing to act on: no benchmark release, no leaderboard, no model to change. If you evaluate reasoning models, the methodological point is the useful one — a strong score on one phrasing of a task says little about an equivalent phrasing, so test the variants before treating a capability as established.

Reasoning models that solve a rule-induction task often fail on a structurally equivalent version of the same task, according to a paper posted to arXiv on 11 September 2026. The authors adapted established rule-induction tasks from cognitive science and generated variants through recombination and substitution — swapping elements while preserving the underlying rule structure. Models gave correct answers on the original forms and then failed the variants. The abstract names no models, reports no scores and gives no evaluation dates; it argues that behaviour cannot be attributed to a general ability, only to the particular contexts tested.

Key facts

  • ·The paper was posted to arXiv on 11 September 2026 as 2609.13948. source
  • ·It extends established rule-induction tasks from cognitive science to test reasoning models. source
  • ·Structurally equivalent variants were created through task isomorphisms, namely recombination and substitution, using each task family's compositional structure. source
  • ·Models that solved a task correctly often failed on structurally equivalent variants of the same task. source
  • ·The abstract reports no named models, no scores and no evaluation dates. source
  • ·The authors conclude many model behaviours lack systematicity, making cognitive abilities hard to establish beyond the contexts evaluated. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire