D2VBench: бенчмарк из 10 000 дилемм проверяет ценностное выравнивание LLM
Команда исследователей выложила на arXiv бенчмарк D2VBench — 10 000 сценариев повседневных дилемм, где сталкиваются несколько ценностей сразу. В основе — 158 вручную размеченных ценностных концепций, а оценка сочетает закрытые и открытые вопросы. Через бенчмарк уже прогнали 8 популярных языковых моделей; датасет открыт на GitHub.
AI-processed from arXiv cs.CL; edited by Hamidun News
Researchers in July 2026 presented on arXiv the D2VBench benchmark — a tool for evaluating the value alignment of large language models across 10000 everyday dilemmas, based on 158 manually labeled value concepts.
What D2VBench Tests
D2VBench evaluates which values lie behind a model's responses when several principles collide at once in an everyday situation — for example, honesty versus care, or personal gain versus the common good. According to the authors, previous test sets poorly covered exactly these multi-factor conflicts and relied on overly simplified evaluation formats.
Each of the 10000 scenarios was assembled in stages — through a collaboration between language models and human annotators. At the core lie 158 fine-grained value concepts, manually labeled, which makes it possible to tie a model's response to a specific value category rather than to general "ethicality."
- Volume: 10000 scenarios of real everyday dilemmas
- Labeling basis: 158 manually annotated value concepts
- Assembly: multi-stage collaboration between LLMs and humans
- Evaluation: a hybrid of closed (multiple-choice) and open-ended questions
- Coverage: 8 common LLMs were tested
- Data: the dataset is published on GitHub (tjunlp-lab/D2VBench)
How the Evaluation Works
Evaluation in D2VBench is built as a hybrid: closed questions with answer options are combined with open-ended questions where the model formulates its answer freely. This format, by the authors' design, captures not only the "correct" choice but also how the model justifies its decision and prioritizes between conflicting values.
The authors ran 8 common language models through the benchmark. According to their data, D2VBench demonstrates high reliability and robustness and reflects models' alignment across different value categories and dimensions — that is, it gives a more detailed picture than a binary "ethical or not" score.
"Existing benchmarks insufficiently cover value dilemmas of everyday
scenarios with multiple value conflicts and use simplified evaluation formats," the D2VBench authors state in their paper on arXiv.
Why It Matters
Value alignment is becoming critical as LLMs enter real everyday scenarios — from health advice to family and workplace conflicts. D2VBench, as noted in the arXiv description, is designed as a more realistic and detailed research tool in this area: it shows not an average "safety" score but a model's behavior along specific value axes.
Publishing the dataset in open access on GitHub means other teams will be able to reproduce the results and run their own models through the 10000 dilemmas, not just the 8 tested by the authors. The breakdown by 158 concepts and two question formats allows models to be compared not by a single number but by a profile of value-based behavior.
What This Means
D2VBench shifts the conversation about AI "safety" from the plane of general declarations into a measurable grid of 158 value concepts. For developers, this is a way to see exactly where a model falters in a value conflict, rather than simply getting a single aggregated score.
Frequently Asked Questions
What is D2VBench?
D2VBench is a benchmark for evaluating the value alignment of large language models, consisting of 10000 everyday dilemma scenarios and based on 158 manually labeled value concepts.
How many scenarios are in D2VBench?
The benchmark contains 10000 scenarios of real everyday dilemmas, each of which was assembled in several stages with the participation of language models and human annotators.
Where can I download the D2VBench dataset?
The dataset is published in open access on GitHub in the tjunlp-lab/D2VBench repository — it can be used to test one's own models.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.