Paper isolates a model's value preferences as an editable direction

single source· 1 articles · confidence: medium · first seen 2026-09-16 20:00 UTC

What this means for you

Nothing to build on yet: this is a preprint, and the abstract does not say whether the dataset, code or weights are released. If you run a small multilingual model, the cheap check worth copying is position bias — Llama-3.2-1B and 3B answered by option order here, and the authors report fine-tuning removed it.

A 16 September arXiv preprint describes a 12,000-item dataset of two-option moral dilemmas covering honesty against justice, justice against autonomy, and autonomy against honesty, with translations into Hindi, Arabic, Spanish and Chinese. The authors report that GPT-5-mini picks honesty over autonomy in all five languages when no policy is given, and that Llama-3.2-1B and 3B strongly favour whichever option is listed first. Fine-tuning, plain or on preferred-versus-rejected pairs, removed that bias and took benchmark accuracy above 98%, with no evaluation date given. A second experiment computes task vectors — weight differences that carry a behaviour — and orthogonalises the value direction against instruction-following to isolate it.

Models in this story

Key facts

  • ·The paper describes a 12,000-instance dataset of two-option dilemmas covering three pairwise value conflicts: honesty vs justice, justice vs autonomy, and autonomy vs honesty. source
  • ·The dilemmas were translated into Hindi, Arabic, Spanish and Chinese, giving five languages in total. source
  • ·GPT-5-mini favoured honesty over autonomy in all five languages when no policy was given. source
  • ·Llama-3.2-1B and 3B showed strong first-option bias; the authors report that plain fine-tuning and Direct Preference Optimization removed it, raising accuracy above 98%. source
  • ·The paper is arXiv preprint 2609.21094, posted 16 September 2026. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire