Fine-tuning reviewers on model-written reviews narrows the ratings they give
single source· 1 articles · confidence: medium · first seen 2026-09-16 20:00 UTC
What this means for you
If you assemble review data for training, provenance now matters: in this study, mixing model-written reviews into the corpus narrowed the ratings the reviewer produced. Nothing to deploy today — the paper gives no evaluation numbers and no link to TrustReviewer, so neither the effect size nor the fix can be checked.
Fine-tuning a reviewer on reviews that other models wrote makes it less varied, a preprint reports. Starting from Llama 3.1 8B, the authors fine-tuned on official ICLR reviews from 2018–2023, then trained four successors on ICLR 2024 data using different mixes of official and model-generated reviews. Adding synthetic reviews compressed rating distributions and reduced semantic diversity within a paper and across the corpus — a pattern the authors call "scientific-judgment collapse". Their proposed fix, TrustReviewer, which they describe as open source, pairs a curated training corpus with activation steering (nudging internal states at reply time rather than retraining). No evaluation numbers or release link.
Models in this story
Key facts
- ·The study starts from Llama 3.1 8B, fine-tuned on official ICLR reviews from 2018 to 2023, and then trains four successor reviewers on ICLR 2024 data. source
- ·The four successors were trained on systematically varied mixtures of official reviews and model-generated reviews. source
- ·Adding synthetic reviews compressed rating distributions and reduced semantic diversity within the same paper and across the corpus, a pattern the authors name "scientific-judgment collapse". source
- ·The proposed mitigation, TrustReviewer, trains the core reviewer in a single stage on a curated corpus and applies paired activation steering at test time, without further training or additional expert annotation. source
- ·The preprint is arXiv 2609.20942, dated 16 September 2026; the available text reports no benchmark scores and no evaluation dates. source
What the sources say
- Hugging Face Daily Papers (research) — Controlled ICLR experiment testing whether reviewers trained on model-written reviews lose judgment variety.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersWhen AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation2026-09-16