Automated judge at 0.36% of the cost comes within three points

single source· 1 articles · confidence: medium · first seen 2026-09-21 20:00 UTC

What this means for you

If you use an LLM to grade outputs at scale, this is the shape of the saving: a cheap first pass on preference and factuality calls, with a second grader only on the ones it flags as uncertain. Don't plan around it yet — the material available says nothing about code, weights or a hosted endpoint, and nothing here has been replicated.

A single arXiv preprint, posted 21 September 2026, reports that a decision-only automated judge — one that returns a verdict rather than a written critique — comes within three percentage points of the strongest of sixteen other judges on ordinary preference and factuality grading, at 0.36% of that comparator's fee. The others were generative models and reward models, which score an answer directly instead of explaining it; adjudication was blind. Gaps widen when a derivation must be checked or a well-written wrong answer rejected. A frozen cascade that takes confident verdicts and escalates the rest kept 99% of the comparator's accuracy.

Key facts

  • ·The preprint is arXiv 2609.26550, posted 21 September 2026. source
  • ·The judge under test, called JEV, is compared against sixteen generative and reward-model judges, with blinded human adjudication. source
  • ·It is reported within three percentage points of its strongest comparator on ordinary preference and evidence-grounded factuality. source
  • ·It is reported to run at 0.36% of that comparator's fee. source
  • ·A frozen cascade that accepts confident verdicts and escalates uncertain ones is reported to retain 99% of the comparator's accuracy. source
  • ·The gap widens on judgements that require checking a derivation or resisting an elaborately written wrong answer, and on several benchmarks is concentrated in low-confidence decisions. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire