Skip to content
Road to Intelligence

Papers

The family tree

Research is a conversation across decades. Each line joins a paper to one it built on. Pick a paper to light up its ancestors (blue) and descendants (orange). Rows follow the chapter where each paper is discussed.
1. What Is Artificial Intelligence?2. Math Toolkit3. Machine Learning4. Neural Networks5. Vision, Speech & Reinforcement Learning6. Language Before Transformers7. Transformers8. Rise of Large Language ModelsBeyond (later chapters)19011943194819501955195819621966196819761980198219861989199219941995200120032006200920102012201520202025Bengio et al. 2003: A Neural Probabilistic Language ModelMikolov et al. 2013: Efficient Estimation of Word Representations in Vector SpaceMikolov et al. 2013: Distributed Representations of Words and Phrases and their CompositionalityPennington et al. 2014: GloVe: Global Vectors for Word RepresentationCho et al. 2014: Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine TranslationSutskever et al. 2014: Sequence to Sequence Learning with Neural NetworksSutskever et al. 2014Bahdanau et al. 2014: Neural Machine Translation by Jointly Learning to Align and TranslateBahdanau et al. 2014Vaswani et al. 2017: Attention Is All You NeedVaswani et al. 2017He et al. 2015: Deep Residual Learning for Image RecognitionHe et al. 2015Ba et al. 2016: Layer NormalizationBa et al. 2016Devlin et al. 2018: BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingDevlin et al. 2018Raffel et al. 2019: Exploring the Limits of Transfer Learning with a Unified Text-to-Text TransformerRaffel et al. 2019Xiong et al. 2020: On Layer Normalization in the Transformer ArchitectureXiong et al. 2020Sennrich et al. 2015: Neural Machine Translation of Rare Words with Subword UnitsBrown et al. 2020: Language Models are Few-Shot LearnersBrown et al. 2020Kaplan et al. 2020: Scaling Laws for Neural Language ModelsKaplan et al. 2020Hoffmann et al. 2022: Training Compute-Optimal Large Language ModelsKingma et al. 2014: Adam: A Method for Stochastic OptimizationShannon 1948: A Mathematical Theory of CommunicationRobbins et al. 1951: A Stochastic Approximation MethodKullback et al. 1951: On Information and SufficiencyRuder 2016: An overview of gradient descent optimization algorithmsLoshchilov et al. 2017: Decoupled Weight Decay RegularizationMcCulloch et al. 1943: A logical calculus of the ideas immanent in nervous activityTuring 1950: Computing Machinery and IntelligenceMcCarthy et al. 1955: A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence, August 31, 1955Rosenblatt 1958: The perceptron: A probabilistic model for information storage and organization in the brain.Samuel 1959: Some Studies in Machine Learning Using the Game of CheckersWeizenbaum 1966: ELIZA—a computer program for the study of natural language communication between man and machineHart et al. 1968: A Formal Basis for the Heuristic Determination of Minimum Cost PathsNewell et al. 1976: Computer science as empirical inquiryRumelhart et al. 1986: Learning representations by back-propagating errorsKrizhevsky et al. 2012: ImageNet Classification with Deep Convolutional Neural NetworksMnih et al. 2015: Human-level control through deep reinforcement learningSilver et al. 2016: Mastering the game of Go with deep neural networks and tree searchChristiano et al. 2017: Deep reinforcement learning from human preferencesRadford et al. 2021: Learning Transferable Visual Models From Natural Language SupervisionRadford et al. 2021Wei et al. 2022: Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsOuyang et al. 2022: Training language models to follow instructions with human feedbackDeepSeek-AI et al. 2025: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningDomingos 2012: A few useful things to know about machine learningKaufman et al. 2012: Leakage in data miningCortes et al. 1995: Support-vector networksBreiman 2001: Random ForestsTibshirani 1996: Regression Shrinkage and Selection Via the LassoPearson 1901: On lines and planes of closest fit to systems of points in spaceLloyd 1982: Least squares quantization in PCMHornik et al. 1989: Multilayer feedforward networks are universal approximatorsCybenko 1989: Approximation by superpositions of a sigmoidal functionBengio et al. 1994: Learning long-term dependencies with gradient descent is difficultHochreiter et al. 1997: Long Short-Term MemoryLeCun et al. 1998: Gradient-based learning applied to document recognitionGlorot et al. 2010: Understanding the difficulty of training deep feedforward neural networksSrivastava et al. 2014: Dropout: A Simple Way to Prevent Neural Networks from OverfittingIoffe et al. 2015: Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate ShiftHe et al. 2015: Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet ClassificationHendrycks et al. 2016: Gaussian Error Linear Units (GELUs)Hubel et al. 1962: Receptive fields, binocular interaction and functional architecture in the cat's visual cortexFukushima 1980: Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in positionLeCun et al. 1989: Backpropagation Applied to Handwritten Zip Code RecognitionDeng et al. 2009: ImageNet: A large-scale hierarchical image databaseRen et al. 2015: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal NetworksRonneberger et al. 2015: U-Net: Convolutional Networks for Biomedical Image SegmentationDosovitskiy et al. 2020: An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleDosovitskiy et al. 2020Graves et al. 2006: Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networksHinton et al. 2012: Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research GroupsOord et al. 2016: WaveNet: A Generative Model for Raw AudioRadford et al. 2022: Robust Speech Recognition via Large-Scale Weak SupervisionRadford et al. 2022Sutton et al. 2018: Reinforcement Learning: An Introduction (2nd edition)Watkins et al. 1992: Q-learningWilliams 1992: Simple statistical gradient-following algorithms for connectionist reinforcement learningMnih et al. 2016: Asynchronous Methods for Deep Reinforcement LearningSchulman et al. 2017: Proximal Policy Optimization AlgorithmsPeters et al. 2018: Deep contextualized word representationsHoward et al. 2018: Universal Language Model Fine-tuning for Text ClassificationRadford et al. 2018: Improving Language Understanding by Generative Pre-TrainingRadford et al. 2018Radford et al. 2019: Language Models are Unsupervised Multitask LearnersFan et al. 2018: Hierarchical Neural Story GenerationHoltzman et al. 2019: The Curious Case of Neural Text DegenerationBommasani et al. 2021: On the Opportunities and Risks of Foundation ModelsWei et al. 2021: Finetuned Language Models Are Zero-Shot LearnersMin et al. 2022: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Olsson et al. 2022: In-context Learning and Induction HeadsWei et al. 2022: Emergent Abilities of Large Language ModelsSchaeffer et al. 2023: Are Emergent Abilities of Large Language Models a Mirage?Delétang et al. 2023: Language Modeling Is CompressionTouvron et al. 2023: LLaMA: Open and Efficient Foundation Language ModelsOpenAI et al. 2023: GPT-4 Technical Report

built on (ancestors) built on it (descendants)Larger dots: essential papers. Click a dot, or pick from the list.

2017 · essential

Attention Is All You Need

Introduced the Transformer — the architecture behind BERT, GPT and nearly every modern large language model, and later adapted to vision, audio and more.

Built directly on

Directly influenced

Links record direct influence as described in each paper and in the chapters; they are not a complete citation graph. Browse every paper with its summary in the paper library.