Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Multimodal Hallucination and Evaluation

Should knowUnderstand12 minDifficulty

Evaluating multimodal models means checking both directions separately: whether a model that reads images reports only what is there, and whether a model that makes images produces realistic, varied, on-prompt and original outputs.

The problem

Fluent captions and beautiful pictures look like success, but a captioner can describe objects that are not in the image and a generator can match its benchmark while ignoring parts of the prompt or copying training images.

The solution

Use targeted probes (yes/no questions about present and absent objects), expert benchmarks with error analysis, distribution metrics such as FID together with alignment metrics such as CLIPScore, and human judgment, each for the failure it can detect.

The consequence

No single number captures multimodal quality; reports should say which failure each metric checks and what it misses.

Reading: does it say only what is there?

A vision-language model's answer comes from a language model that is very good at plausible text. Li and colleagues found that large vision-language models mostly suffer from severe object hallucination, and that objects frequent in the visual instruction data, or that often co-occur with the objects actually in the image, are especially likely to be hallucinated. Established A kitchen photo invites "a refrigerator" whether or not one is visible.

Their POPE method asks yes/no questions about objects, with absent objects drawn three ways: at random, from the most frequent objects, and adversarially from objects that often co-occur with what is present. Established The adversarial setting is the informative one, because it asks exactly where the model's priors and the pixels disagree.

For harder questions, MMMU collected 11.5K college-level questions across 30 subjects and 30 image types (charts, diagrams, chemical structures, sheet music); at release GPT-4V scored 56% and Gemini Ultra 59%, and in 150 analysed GPT-4V errors, 35% were perceptual. Established Error analysis like this separates "did not see it" from "saw it but reasoned wrongly", which call for different fixes.

Making: is it good, varied, on-prompt and new?

A generated image set can fail in four separate ways, and metrics differ in which they catch:

QuestionTypical checkMisses
Realistic?FID against real imagesprompt alignment; per-image errors
Varied?precision and recall of features, FIDwhether the variety is on-prompt
On-prompt?CLIPScore, human ratingsCLIP's blind spots (counting, word order)
Original?nearest-neighbour search over training dataanything not in the searched set

FID, introduced by Heusel and colleagues, compares statistics of Inception-network features of generated and real images. Established It is computed over whole sets, so it cannot tell you whether any single image matches its prompt. CLIPScore rates image–text compatibility with CLIP and correlated well with human judgments of captions, but was weaker where captions need outside context, such as news. Established Because CLIP-style models often ignore word order (CLIP), a CLIP-based score can rate "a red cube on a blue sphere" highly for the swapped scene.

Carlini and colleagues extracted over a thousand training images from diffusion models, so originality needs checking too. Established

Tiny example (lab numbers)

The diffusion lab shows why one metric is not enough. At guidance 8 with the prompt "plus", every sample that lands on a shape is on the plus: perfect prompt adherence. Yet fewer than half the samples land on any shape, and only 17% of the plus is covered. A metric that only asks "is it on-prompt?" would call this the best setting in the lab.

A checklist

  1. Which failure does each reported metric detect, and which does it miss?
  2. Are reading results broken down into perception and reasoning errors?
  3. For absent-object probes, are the negatives adversarial or random?
  4. For generation, are quality, diversity, alignment and originality all reported?
  5. Could the benchmark images, or close copies, be in the training data (contamination)?

Chapter 16 takes these questions further: evaluation, reliability and interpretability for every kind of model.

Mini experiment

In the diffusion lab, choose "plus" and compare guidance 3 and guidance 5. Their "On the plus" percentages are close; their coverage is not. Which would a prompt-alignment metric prefer, which would a diversity metric prefer, and which would you ship?

What to remember

  • Object hallucination: describing objects that are not in the image; objects common in the training data or co-occurring with what is there are most at risk.
  • POPE asks 'Is there a ⟨object⟩ in the image?' with random, popular and adversarial absent objects.
  • MMMU: 11.5K college-level questions over 30 image types; at release GPT-4V scored 56%, and 35% of its analysed errors were perceptual.
  • FID compares feature statistics of generated and real image sets; it does not check the prompt.
  • CLIPScore checks image–text agreement with CLIP, inheriting CLIP's blind spots.
  • Generators can reproduce training images; check originality as well as quality.

Key papers

Optional

GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Martin Heusel, Hubert Ramsauer et al. · 2017

Introduced the Fréchet Inception Distance (FID), still the most reported number for image-generation quality, so it is worth knowing what it measures and what it misses.

How to read it: Most readers only need the FID definition. Note that it compares feature statistics of whole sets: it says nothing about whether one image matches its prompt.

~40 min readarXiv:1706.08500✓ verified 2026-10-07
Important

When and why vision-language models behave like bags-of-words, and what to do about it?

Mert Yuksekgonul, Federico Bianchi et al. · 2022

Measured, at scale, that contrastive image–text models often ignore word order and mis-bind attributes, and explained why their training and benchmarks let them.

How to read it: The key argument is about the training objective: if shuffled captions still retrieve the right image, the contrastive loss never needed word order. Hard negatives that differ only in order fix the incentive.

~35 min readarXiv:2210.01936✓ verified 2026-10-07
Optional

CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Jack Hessel, Ari Holtzman et al. · 2021

Using CLIP's image–text similarity as a metric became standard for captioning and text-to-image evaluation, along with its blind spots.

How to read it: Read the case studies at the end: a metric built on one model inherits that model's blind spots.

~25 min readarXiv:2104.08718✓ verified 2026-10-07
Important

Extracting Training Data from Diffusion Models

Nicholas Carlini, Jamie Hayes et al. · 2023

Showed that image generators can reproduce individual training images, which matters for privacy and copyright and for what 'generating new images' means.

How to read it: Look at how 'memorized' is defined before reading the counts, and note the role of duplicated training images.

~40 min readarXiv:2301.13188✓ verified 2026-10-07
Important

Evaluating Object Hallucination in Large Vision-Language Models

Yifan Li, Yifan Du et al. · 2023

A systematic look at vision-language models describing objects that are not in the image, and POPE, a simple yes/no probe for it.

How to read it: The three negative-sampling settings are the clever part: an adversarial absent object is one that often appears alongside what is really there.

~30 min readarXiv:2305.10355✓ verified 2026-10-07
Optional

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Xiang Yue, Yuansheng Ni et al. · 2023

A widely reported benchmark of college-level questions that need an image (charts, diagrams, chemical structures, sheet music) and subject knowledge.

How to read it: Read the error analysis: how many failures are perception errors, how many knowledge errors, how many reasoning errors.

~30 min readarXiv:2311.16502✓ verified 2026-10-07