arxiv:2609.21094
Published on Sep 17
· Submitted by
Utkarsh Agarwal on Sep 21
· Mohamed Bin Zayed University of Artificial Intelligence
Upvote
3
Authors:
Abstract
Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even strong models exhibit hidden biases and brittle instruction-following across languages. We introduce a 12,000-instance dataset of two-option dilemmas covering pairwise three value conflicts: Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, along with their translations into Hindi, Arabic, Spanish, and Chinese, to probe cross-lingual behavior. Benchmarking on GPT-5-mini reveals that it consistently favors Honesty over Autonomy across all five languages when no policy is given. The Llama-3.2-1/3B models exhibit strong first-option bias; however, both plain fine-tuning and Direct Preference Optimization fine-tuning effectively remove this bias, increasing accuracy to greater than 98%. In order to decouple the effect of learning correlations in the dataset from abstract values, we propose a task vector transfer based experiment where after computing the task vectors for a direction of value preference we orthogonalize it with respect to the general instruction following vector. Our experiment shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.
View arXiv page View PDF GitHub 1 Add to collection
Community
Paper author Paper submitter 1 day ago
https://x.com/utkarshag0203/status/2075791098937823623?s=20
about 13 hours ago
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
-
Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales (2026)
-
Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization (2026)
-
Learning When to Trust via Selective Context Preference Optimization (2026)
-
Language Chain in Alignment: Cross-lingual Ranking Preference Optimization (2026)
-
AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection (2026)
-
CoCoA: Context-Conditional Cultural Alignment for Large Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on HF Mirror checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
· Sign up or log in to comment
Upvote
3
Get this paper in your agent:
hf papers read 2609.21094
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.21094 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.21094 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.21094 in a Space README.md to link it from this page.