Fine-tuning reviewers on model-written reviews narrows the ratings they give

single source· 1 articles · confidence: medium · first seen 2026-09-16 20:00 UTC

What this means for you

If you assemble review data for training, provenance now matters: in this study, mixing model-written reviews into the corpus narrowed the ratings the reviewer produced. Nothing to deploy today — the paper gives no evaluation numbers and no link to TrustReviewer, so neither the effect size nor the fix can be checked.

Fine-tuning a reviewer on reviews that other models wrote makes it less varied, a preprint reports. Starting from Llama 3.1 8B, the authors fine-tuned on official ICLR reviews from 2018–2023, then trained four successors on ICLR 2024 data using different mixes of official and model-generated reviews. Adding synthetic reviews compressed rating distributions and reduced semantic diversity within a paper and across the corpus — a pattern the authors call "scientific-judgment collapse". Their proposed fix, TrustReviewer, which they describe as open source, pairs a curated training corpus with activation steering (nudging internal states at reply time rather than retraining). No evaluation numbers or release link.

Models in this story

Key facts

  • ·The study starts from Llama 3.1 8B, fine-tuned on official ICLR reviews from 2018 to 2023, and then trains four successor reviewers on ICLR 2024 data. source
  • ·The four successors were trained on systematically varied mixtures of official reviews and model-generated reviews. source
  • ·Adding synthetic reviews compressed rating distributions and reduced semantic diversity within the same paper and across the corpus, a pattern the authors name "scientific-judgment collapse". source
  • ·The proposed mitigation, TrustReviewer, trains the core reviewer in a single stage on a curated corpus and applies paired activation steering at test time, without further training or additional expert annotation. source
  • ·The preprint is arXiv 2609.20942, dated 16 September 2026; the available text reports no benchmark scores and no evaluation dates. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire