Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Reading Documents, Charts and Screens

Should knowUnderstand10 minDifficulty

Document understanding means extracting text, numbers and structure from images of pages, charts, forms and screens, either with a separate OCR step or, increasingly, by a vision-language model reading the pixels directly.

The problem

Invoices, scanned contracts, dashboards and screenshots hold information as pixels; a text model cannot read them, and a generic image model blurs exactly the small characters that matter.

The solution

Run OCR and feed the text (plus layout) to a language model, or train an end-to-end model that reads text from the image at a resolution high enough to keep the characters.

The consequence

Charts, receipts and screenshots become queryable, but numbers can be misread or invented, and end-to-end models give no OCR stage to inspect when they are wrong.

Two architectures

OCR pipeline. A specialised model finds and reads the text, returning words with bounding boxes; a language model answers from that text. Every stage can be inspected: you can see exactly which characters OCR returned.

End to end. A vision-language model reads the page image itself. Donut was proposed as an OCR-free document-understanding Transformer, motivated by the cost of OCR engines, their inflexibility across languages and document types, and OCR errors propagating to later stages. Established Most current multimodal assistants read documents this way.

The trade-off is the familiar one between a pipeline and a monolith: the pipeline is debuggable stage by stage, the monolith can use layout and context together but fails without telling you which part failed. Interpretation

Resolution is everything

A character 3 pixels wide and 5 high occupies a sliver of a 14- or 16-pixel patch. If the image is downscaled before encoding, or a codebook rounds the patch to a generic "grey texture", the characters are gone before the language model sees anything. The tokens lab shows this directly: in every setting, the error inside the receipt's or chart's text is higher than elsewhere in the picture. Systems that handle documents well use higher input resolution or cut the page into tiles, which multiplies the visual token count.

Charts are reading plus reasoning

ChartQA collected 9.6K human-written questions and 23.1K generated questions about charts, many of which need several logical and arithmetic operations and refer to visual features of the chart. Established "Which month dipped, and by how much?" requires reading two bar heights, subtracting, and naming the month: perception errors and reasoning errors compound. MMMU's analysis of 150 GPT-4V errors attributed 35% to perceptual errors. Established

Practical checks

For a data engineer, document extraction is an ETL source with a noisy parser:

  • Ask for structured output (JSON with fields) rather than prose.
  • Validate: line items sum to the total, dates parse, IDs match a known pattern.
  • Keep the source crop next to each extracted value so a person can check it.
  • Track error rates per field on a labelled sample, as you would for any parser.

Mini experiment

In the tokens lab, pick the receipt and find the cheapest setting (fewest bits) at which the error inside the text is below 5%. Then compare with the cheapest setting at which the error outside the text is below 5%. What does the gap tell you about how much resolution document reading needs compared with recognizing what kind of image it is?

What to remember

  • Pipeline: OCR → text with positions → language model. Inspectable, but OCR errors propagate.
  • End to end: the VLM reads pixels directly (Donut was an early OCR-free model). Simpler, harder to debug.
  • Small text needs pixels: resolution and patch size decide whether a character survives encoding.
  • Chart questions combine reading values with arithmetic; ChartQA has 9.6K human-written questions.
  • For extraction, ask for structured output and validate it (totals add up, dates parse).

Key papers

Optional

OCR-free Document Understanding Transformer

Geewook Kim, Teakgyu Hong et al. · 2021

Donut read documents straight from pixels instead of running a separate OCR engine first, the design most vision-language models now follow for documents.

How to read it: Compare its error sources with an OCR pipeline's: a pipeline can tell you which stage failed; an end-to-end model cannot.

~30 min readarXiv:2111.15664✓ verified 2026-10-07
Optional

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Xiang Yue, Yuansheng Ni et al. · 2023

A widely reported benchmark of college-level questions that need an image (charts, diagrams, chemical structures, sheet music) and subject knowledge.

How to read it: Read the error analysis: how many failures are perception errors, how many knowledge errors, how many reasoning errors.

~30 min readarXiv:2311.16502✓ verified 2026-10-07