Paper isolates a model's value preferences as an editable direction
single source· 1 articles · confidence: medium · first seen 2026-09-16 20:00 UTC
What this means for you
Nothing to build on yet: this is a preprint, and the abstract does not say whether the dataset, code or weights are released. If you run a small multilingual model, the cheap check worth copying is position bias — Llama-3.2-1B and 3B answered by option order here, and the authors report fine-tuning removed it.
A 16 September arXiv preprint describes a 12,000-item dataset of two-option moral dilemmas covering honesty against justice, justice against autonomy, and autonomy against honesty, with translations into Hindi, Arabic, Spanish and Chinese. The authors report that GPT-5-mini picks honesty over autonomy in all five languages when no policy is given, and that Llama-3.2-1B and 3B strongly favour whichever option is listed first. Fine-tuning, plain or on preferred-versus-rejected pairs, removed that bias and took benchmark accuracy above 98%, with no evaluation date given. A second experiment computes task vectors — weight differences that carry a behaviour — and orthogonalises the value direction against instruction-following to isolate it.
Models in this story
Key facts
- ·The paper describes a 12,000-instance dataset of two-option dilemmas covering three pairwise value conflicts: honesty vs justice, justice vs autonomy, and autonomy vs honesty. source
- ·The dilemmas were translated into Hindi, Arabic, Spanish and Chinese, giving five languages in total. source
- ·GPT-5-mini favoured honesty over autonomy in all five languages when no policy was given. source
- ·Llama-3.2-1B and 3B showed strong first-option bias; the authors report that plain fine-tuning and Direct Preference Optimization removed it, raising accuracy above 98%. source
- ·The paper is arXiv preprint 2609.21094, posted 16 September 2026. source
What the sources say
- Hugging Face Daily Papers (research) — Single preprint abstract: dilemma dataset, cross-lingual bias check, and a weight-editing method for flipping preferences.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersGeometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models2026-09-16