Paper argues calibration should sit beside accuracy in model evaluations
single source· 1 articles · confidence: medium · first seen 2026-09-21 20:00 UTC
What this means for you
Nothing to install — this is a reporting argument, not a metric or a tool. But if your pipeline leans on confidence scores (LLM-as-a-judge, synthetic data, active learning), the paper's point applies: most benchmarks, it says, already record a confidence score and a correctness judgement. Nobody reports them.
A position paper posted to arXiv on 21 September argues that language model evaluations should report calibration — how closely a model's expressed or implied confidence matches whether it is actually right — alongside the accuracy each benchmark already tracks. Standard calibration metrics need only two inputs per example, a confidence score and a correctness judgement, and the authors say most benchmarks already record both, so the number could be reported now. The paper asks each subfield to pair its main metric with a calibration score, and says the two inputs remain undefined for open-ended generation.
Key facts
- ·The paper is posted on arXiv as 2609.26489, dated 21 September 2026, and argues calibration should be treated as an essential property of every model rather than a niche topic. source
- ·Standard calibration metrics require two inputs per example: a confidence score and a correctness judgement. source
- ·The authors state that most benchmarks in use today already provide both inputs, so calibration can be reported immediately. source
- ·For open-ended generation, the paper says defining those two inputs remains an open challenge. source
- ·The paper names LLM-as-a-judge, synthetic data generation and active learning as research methods that assume calibrated confidence without verifying it. source
What the sources say
- Hugging Face Daily Papers (research) — Position paper calling for confidence-versus-correctness reporting alongside accuracy, noting the required inputs are mostly already collected.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersCalibration as a First-Class Criterion in LLM Evaluation2026-09-21