Probe flags answers a model recognises but does not report

single source· 1 articles · confidence: medium · first seen 2026-09-17 20:00 UTC

What this means for you

Nothing to build on yet — no code or weights are mentioned and this is one unreviewed preprint. The useful claim is that an audit can separate a model that will not answer from one that cannot, which is what sandbagging and unlearning checks need. Wait for replication before putting it in a deployment gate.

A preprint describes a way to tell whether a language model is withholding an answer it holds or does not have one. Probe of Internal Recognition shows the model a question and its candidate answers, then reads its internal states to see which candidate it recognises — a forensic Concealed Information Test run inside the network. It needs no reference model or truth corpus. Across eight models from five families it recovered the recognised answer at 0.70–0.87 balanced accuracy. Under prompted deception, trained sandbagging, password-locked and circuit-broken checkpoints it held at 0.85–0.93; after unlearning, it fell back to the level of a never-known question. Not yet peer reviewed.

Key facts

  • ·Probe of Internal Recognition recovered the recognised answer at 0.70–0.87 balanced accuracy across eight models, against a 0.25 chance rate and a 0.28–0.40 unknown-item baseline. source
  • ·Recognition stayed between 0.85 and 0.93 under prompted deception, trained sandbagging, password-locked checkpoints and circuit-broken checkpoints. source
  • ·The method is reference-free: it requires no honest reference model and no labelled truth corpus. source
  • ·When unlearning removed the knowledge, recognition dropped to the level of a question the model never knew. source
  • ·The eight models tested come from five families: Gemma, Qwen, Llama, Mistral and Phi. source
  • ·The work is posted as arXiv 2609.21996, dated 17 September 2026, and has not been peer reviewed. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire