One call spots ten kinds of model misbehaviour, beating trained classifiers

single source· 1 articles · confidence: medium · first seen 2026-09-23 20:00 UTC

What this means for you

If you score model outputs with a large model as judge, this is a cheaper option worth testing: one call, no task-specific training, code and data published. But the paper's finding is that what the detector is shown matters more than how it is asked, mainly through fields carrying the label — check that leak before relying on the 0.886.

A paper posted to arXiv on 23 September tests Jev, a model trained to answer many typed questions about one input in a single call with calibrated probabilities, as a detector of ten alignment failures, from sycophancy and jailbreaks to hallucination and reward hacking. Across 44 benchmarks and five target models, one generic question reached a median AUROC (a separation score; 0.5 is chance) of 0.886 without task-specific training, beating supervised baselines on most. It matched the reference scorer's agreement with human labels on the two benchmarks scored by people, and costs 63 times less than LLM-judge scorers. The paper gives no evaluation date for the score.

Key facts

  • ·RLCDAlignBench spans ten alignment failure types, 44 benchmarks and five target models. source
  • ·One generic question reached a median AUROC of 0.886 zero-shot, with no task-specific training; the paper gives no evaluation date for the score. source
  • ·The method beats supervised baselines on most of the benchmarks tested. source
  • ·It costs 63 times less than LLM-judge scorers, according to the paper. source
  • ·It matched the reference scorer's agreement with human labels on the two benchmarks whose labels came from people. source
  • ·The paper reports that question wording matters little, while the context matters more, mostly through fields that encode the label. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire