Concept · Chapter 15: Multimodal AI
Reading Documents, Charts and Screens
Document understanding means extracting text, numbers and structure from images of pages, charts, forms and screens, either with a separate OCR step or, increasingly, by a vision-language model reading the pixels directly.
The problem
Invoices, scanned contracts, dashboards and screenshots hold information as pixels; a text model cannot read them, and a generic image model blurs exactly the small characters that matter.
The solution
Run OCR and feed the text (plus layout) to a language model, or train an end-to-end model that reads text from the image at a resolution high enough to keep the characters.
The consequence
Charts, receipts and screenshots become queryable, but numbers can be misread or invented, and end-to-end models give no OCR stage to inspect when they are wrong.
You should understand first
- Vectors
- Dot Product
- Embeddings
- Attention
- Probability and Distributions
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Text Embeddings
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Contrastive Learning
- Text as Data
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Vision Transformer (ViT)
- Audio and Spectrograms
- Turning Signals into Tokens
- CLIP and Image–Text Embeddings
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Pretrain, Then Fine-Tune
- Supervised Fine-Tuning
- Vision-Language Models
- Reading Documents, Charts and Screens
Two architectures
OCR pipeline. A specialised model finds and reads the text, returning words with bounding boxes; a language model answers from that text. Every stage can be inspected: you can see exactly which characters OCR returned.
End to end. A vision-language model reads the page image itself. Donut was proposed as an OCR-free document-understanding Transformer, motivated by the cost of OCR engines, their inflexibility across languages and document types, and OCR errors propagating to later stages. Established Most current multimodal assistants read documents this way.
The trade-off is the familiar one between a pipeline and a monolith: the pipeline is debuggable stage by stage, the monolith can use layout and context together but fails without telling you which part failed. InterpretationResolution is everything
A character 3 pixels wide and 5 high occupies a sliver of a 14- or 16-pixel patch. If the image is downscaled before encoding, or a codebook rounds the patch to a generic "grey texture", the characters are gone before the language model sees anything. The tokens lab shows this directly: in every setting, the error inside the receipt's or chart's text is higher than elsewhere in the picture. Systems that handle documents well use higher input resolution or cut the page into tiles, which multiplies the visual token count.
Charts are reading plus reasoning
ChartQA collected 9.6K human-written questions and 23.1K generated questions about charts, many of which need several logical and arithmetic operations and refer to visual features of the chart. Established "Which month dipped, and by how much?" requires reading two bar heights, subtracting, and naming the month: perception errors and reasoning errors compound. MMMU's analysis of 150 GPT-4V errors attributed 35% to perceptual errors. Established
Practical checks
For a data engineer, document extraction is an ETL source with a noisy parser:
- Ask for structured output (JSON with fields) rather than prose.
- Validate: line items sum to the total, dates parse, IDs match a known pattern.
- Keep the source crop next to each extracted value so a person can check it.
- Track error rates per field on a labelled sample, as you would for any parser.
Mini experiment
In the tokens lab, pick the receipt and find the cheapest setting (fewest bits) at which the error inside the text is below 5%. Then compare with the cheapest setting at which the error outside the text is below 5%. What does the gap tell you about how much resolution document reading needs compared with recognizing what kind of image it is?
What to remember
- Pipeline: OCR → text with positions → language model. Inspectable, but OCR errors propagate.
- End to end: the VLM reads pixels directly (Donut was an early OCR-free model). Simpler, harder to debug.
- Small text needs pixels: resolution and patch size decide whether a character survives encoding.
- Chart questions combine reading values with arithmetic; ChartQA has 9.6K human-written questions.
- For extraction, ask for structured output and validate it (totals add up, dates parse).
Key papers
OCR-free Document Understanding Transformer
Geewook Kim, Teakgyu Hong et al. · 2021
Donut read documents straight from pixels instead of running a separate OCR engine first, the design most vision-language models now follow for documents.
How to read it: Compare its error sources with an OCR pipeline's: a pipeline can tell you which stage failed; an end-to-end model cannot.
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Xiang Yue, Yuansheng Ni et al. · 2023
A widely reported benchmark of college-level questions that need an image (charts, diagrams, chemical structures, sheet music) and subject knowledge.
How to read it: Read the error analysis: how many failures are perception errors, how many knowledge errors, how many reasoning errors.