Apple ML Research→ original

Apple ML Research изучила alignment в мультимодальных LLM и их склонность к галлюцинациям

Apple ML Research опубликовала исследование о preference alignment в мультимодальных LLM. В отличие от текстовых моделей, MLLM галлюцинируют двояко: называют неверные факты или описывают изображения неточно. Цель alignment — привязать ответы модели к реальному содержанию картинки. Авторы называют это направление критически важным, но до сих пор малоизученным.

AI-processed from Apple ML Research; edited by Hamidun News
Apple ML Research изучила alignment в мультимодальных LLM и их склонность к галлюцинациям
Source: Apple ML Research. Collage: Hamidun News.
◐ Listen to article

Apple ML Research has published a comprehensive study on preference alignment in multimodal language models (MLLMs) — systems capable of simultaneously processing text and images. The authors argue that alignment, which has become the standard for text-based LLMs, remains significantly less studied in the multimodal context.

Why MLLM Hallucinations Are More Complex Than Ordinary Ones

Multimodal models hallucinate differently from text-based LLMs. According to the Apple ML Research study, in image-understanding systems hallucinations manifest in two forms: factual errors — when the model states incorrect facts about the world — and content inconsistency — when the response sounds convincing but contradicts what is actually depicted in the image.

The second type is particularly insidious: if a model describes a blue car in a photo as red, or labels a summer image as winter — this is not an error in knowledge, but an error in "vision." It is precisely against this that alignment for multimodal systems is directed.

"The primary goal of alignment for MLLMs is to encourage these models to ground their responses in the image information," states the

Apple ML Research publication.

How Preference Alignment Works

Preference alignment is a fine-tuning technique in which the model is presented with pairs of responses and labeled as to which is "preferable" according to given criteria. In text-based LLMs, this approach — including RLHF (Reinforcement Learning from Human Feedback) and its derivatives — has become the de facto standard since the advent of InstructGPT and ChatGPT. For multimodal systems, almost no such systematic studies had been conducted until now.

Key theses of the paper:

  • Preference alignment is critically important for LLM quality, but its impact in MLLMs is barely studied
  • In MLLMs, alignment addresses an additional challenge: grounding the model's response in real visual content
  • Hallucinations in MLLMs are divided into two types: factual and visual (inconsistency with the image)
  • Apple ML Research conducts a systematic comparative analysis of alignment techniques applied to image-understanding tasks

According to the study, this work is one of the first comprehensive attempts to bridge the gap between how well alignment is understood for text-based models and how little is known about its impact on multimodal systems.

Why This Matters for the Industry

Multimodal models are now used in medical diagnostics, content moderation, and consumer assistants — from Google Lens to Apple Intelligence. In each of these scenarios, the hallucination of "seeing something that isn't there" costs incomparably more than a textual error: an incorrectly read X-ray or a misidentified object on the road carries direct consequences.

The Apple ML Research study arrives at a moment when multimodal systems are transitioning from laboratory prototypes to mass deployment. Understanding how alignment shapes the accuracy of a model's "vision" is becoming a fundamental question for the entire industry.

What This Means

The Apple ML Research paper identifies a systemic gap: the industry is actively developing and deploying multimodal models, but the mechanism for aligning them with the actual content of images has remained poorly understood. If the study's findings form the basis for new training methods, this will directly affect the quality of products that work with visual content — including Apple Intelligence and similar systems.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…