arXiv cs.AI→ original

Бенчмарк Learn2Discern: LLM почти не отличают надёжные источники от популярных

Исследователи представили бенчмарк Learn2Discern и проверили на нём 13 языковых моделей — почти 670 000 тестов. Итог: модели отличают надёжные источники от ненадёжных на уровне случайного угадывания и опираются на популярность источника вдвое чаще, чем на его достоверность. Опрос 299 пользователей подтвердил: такие сбои подрывают доверие к ИИ.

AI-processed from arXiv cs.AI; edited by Hamidun News
Бенчмарк Learn2Discern: LLM почти не отличают надёжные источники от популярных
Source: arXiv cs.AI. Collage: Hamidun News.
◐ Listen to article

In July 2026, researchers published a paper on arXiv about "information discernment" in large language models and introduced the Learn2Discern (L2D) benchmark. Across 13 models and almost 670,000 trials, the finding was clear: models distinguish reliable sources from unreliable ones at roughly the level of random guessing.

What Learn2Discern tests

Learn2Discern evaluates two skills a model needs when working with external data: source discernment — whether the model trusts a reliable source more strongly, and truth discernment — whether it updates its answer more strongly when a new claim brings it closer to the truth. The framework is built on three normative axioms with interpretable metrics. According to the authors, real users share these principles: a pre-registered survey with a quota sample (n=299) showed that violating them lowers trust in the model and willingness to use it.

Key findings

On both dimensions, the models fail: both source reliability and closeness to the truth land at roughly random levels.

  • 13 models, almost 670,000 trials
  • Models rely on a source's popularity roughly twice as strongly as on its reliability
  • Answers are updated about equally regardless of whether a claim moves closer to or further from the truth
  • Models integrate external knowledge best in cases where their own priors are already the most accurate
  • User survey: n=299, quota sample, pre-registered
"Models perform close to random guessing on both source reliability

and truthfulness," the study's abstract on arXiv states.

Why newer models don't help

Larger, newer models improve only truth discernment, not source discernment. This is a blind spot that isn't fixed by increasing model complexity: the ability to distinguish a credible source from a popular one does not grow with size. Worse, according to Learn2Discern, models integrate external knowledge most effectively exactly where their own priors are already the most accurate — in other words, they help themselves where help is needed least. On a positive note, the authors did find that simple inference-time interventions improve both skills at once, without any model fine-tuning.

What this means

As LLMs replace traditional search, the ability to tell a credible source apart from a merely popular one is becoming a critical alignment property. So far, judging by the 13 models tested, this skill hasn't been learned — and model size doesn't solve the problem. The authors have released the dataset and survey as an open test bed for this task.

FAQ

What is Learn2Discern?

Learn2Discern (L2D) is an experimental framework and benchmark for evaluating how LLMs weigh external information: whether they trust reliable sources and whether they move closer to the truth. It is built on three normative axioms with interpretable metrics; the authors have made the dataset and survey open for reuse.

Which models were tested?

The study covered 13 models across almost 670,000 trials. All of them showed source and truth discernment close to random levels, while newer, larger models improved only truth discernment, not source discernment.

Why does this matter for AI search?

As LLMs replace search engines, they decide on their own which source to trust. If a model relies on a source's popularity roughly twice as strongly as on its reliability, it risks amplifying widespread but incorrect claims.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…