One call spots ten kinds of model misbehaviour, beating trained classifiers
single source· 1 articles · confidence: medium · first seen 2026-09-23 20:00 UTC
What this means for you
If you score model outputs with a large model as judge, this is a cheaper option worth testing: one call, no task-specific training, code and data published. But the paper's finding is that what the detector is shown matters more than how it is asked, mainly through fields carrying the label — check that leak before relying on the 0.886.
A paper posted to arXiv on 23 September tests Jev, a model trained to answer many typed questions about one input in a single call with calibrated probabilities, as a detector of ten alignment failures, from sycophancy and jailbreaks to hallucination and reward hacking. Across 44 benchmarks and five target models, one generic question reached a median AUROC (a separation score; 0.5 is chance) of 0.886 without task-specific training, beating supervised baselines on most. It matched the reference scorer's agreement with human labels on the two benchmarks scored by people, and costs 63 times less than LLM-judge scorers. The paper gives no evaluation date for the score.
Key facts
- ·RLCDAlignBench spans ten alignment failure types, 44 benchmarks and five target models. source
- ·One generic question reached a median AUROC of 0.886 zero-shot, with no task-specific training; the paper gives no evaluation date for the score. source
- ·The method beats supervised baselines on most of the benchmarks tested. source
- ·It costs 63 times less than LLM-judge scorers, according to the paper. source
- ·It matched the reference scorer's agreement with human labels on the two benchmarks whose labels came from people. source
- ·The paper reports that question wording matters little, while the context matters more, mostly through fields that encode the label. source
What the sources say
- Hugging Face Daily Papers (research) — Introduces the benchmark and reports that a single generic question separates ten failure types better than trained baselines.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures2026-09-23