Paper1901 · Pearson 1901 · On lines and planes of closest fit to systems of points in space 1943 · McCulloch et al. 1943 · A logical calculus of the ideas immanent in nervous activity 1948 · Shannon 1948 · A Mathematical Theory of Communication 1950 · Turing 1950 · Computing Machinery and Intelligence 1951 · Kullback et al. 1951 · On Information and Sufficiency 1951 · Robbins et al. 1951 · A Stochastic Approximation Method 1955 · McCarthy et al. 1955 · A Proposal for the Dartmouth Summer Research Project on Artificial I… 1958 · Rosenblatt 1958 · The perceptron: A probabilistic model for information storage and or… 1959 · Samuel 1959 · Some Studies in Machine Learning Using the Game of Checkers 1962 · Hubel et al. 1962 · Receptive fields, binocular interaction and functional architecture … 1966 · Weizenbaum 1966 · ELIZA—a computer program for the study of natural language communica… 1968 · Hart et al. 1968 · A Formal Basis for the Heuristic Determination of Minimum Cost Paths 1976 · Newell et al. 1976 · Computer science as empirical inquiry 1980 · Fukushima 1980 · Neocognitron: A self-organizing neural network model for a mechanism… 1982 · Lloyd 1982 · Least squares quantization in PCM 1986 · Rumelhart et al. 1986 · Learning representations by back-propagating errors 1989 · Cybenko 1989 · Approximation by superpositions of a sigmoidal function 1989 · Hornik et al. 1989 · Multilayer feedforward networks are universal approximators 1989 · LeCun et al. 1989 · Backpropagation Applied to Handwritten Zip Code Recognition 1992 · Watkins et al. 1992 · Q-learning 1992 · Williams 1992 · Simple statistical gradient-following algorithms for connectionist r… 1994 · Bengio et al. 1994 · Learning long-term dependencies with gradient descent is difficult 1995 · Cortes et al. 1995 · Support-vector networks 1996 · Tibshirani 1996 · Regression Shrinkage and Selection Via the Lasso 1997 · Hochreiter et al. 1997 · Long Short-Term Memory 1998 · LeCun et al. 1998 · Gradient-based learning applied to document recognition 2001 · Breiman 2001 · Random Forests 2003 · Bengio et al. 2003 · A Neural Probabilistic Language Model 2006 · Graves et al. 2006 · Connectionist temporal classification: labelling unsegmented sequenc… 2009 · Deng et al. 2009 · ImageNet: A large-scale hierarchical image database 2010 · Glorot et al. 2010 · Understanding the difficulty of training deep feedforward neural net… 2012 · Domingos 2012 · A few useful things to know about machine learning 2012 · Hinton et al. 2012 · Deep Neural Networks for Acoustic Modeling in Speech Recognition: Th… 2012 · Kaufman et al. 2012 · Leakage in data mining 2012 · Krizhevsky et al. 2012 · ImageNet Classification with Deep Convolutional Neural Networks 2013 · Mikolov et al. 2013 · Efficient Estimation of Word Representations in Vector Space 2013 · Mikolov et al. 2013 · Distributed Representations of Words and Phrases and their Compositi… 2014 · Bahdanau et al. 2014 · Neural Machine Translation by Jointly Learning to Align and Translate 2014 · Cho et al. 2014 · Learning Phrase Representations using RNN Encoder-Decoder for Statis… 2014 · Kingma et al. 2014 · Adam: A Method for Stochastic Optimization 2014 · Pennington et al. 2014 · GloVe: Global Vectors for Word Representation 2014 · Srivastava et al. 2014 · Dropout: A Simple Way to Prevent Neural Networks from Overfitting 2014 · Sutskever et al. 2014 · Sequence to Sequence Learning with Neural Networks 2015 · He et al. 2015 · Deep Residual Learning for Image Recognition 2015 · He et al. 2015 · Delving Deep into Rectifiers: Surpassing Human-Level Performance on … 2015 · Ioffe et al. 2015 · Batch Normalization: Accelerating Deep Network Training by Reducing … 2015 · Mnih et al. 2015 · Human-level control through deep reinforcement learning 2015 · Ren et al. 2015 · Faster R-CNN: Towards Real-Time Object Detection with Region Proposa… 2015 · Ronneberger et al. 2015 · U-Net: Convolutional Networks for Biomedical Image Segmentation 2015 · Sennrich et al. 2015 · Neural Machine Translation of Rare Words with Subword Units 2016 · Ba et al. 2016 · Layer Normalization 2016 · Hendrycks et al. 2016 · Gaussian Error Linear Units (GELUs) 2016 · Mnih et al. 2016 · Asynchronous Methods for Deep Reinforcement Learning 2016 · Oord et al. 2016 · WaveNet: A Generative Model for Raw Audio 2016 · Ruder 2016 · An overview of gradient descent optimization algorithms 2016 · Silver et al. 2016 · Mastering the game of Go with deep neural networks and tree search 2017 · Christiano et al. 2017 · Deep reinforcement learning from human preferences 2017 · Loshchilov et al. 2017 · Decoupled Weight Decay Regularization 2017 · Schulman et al. 2017 · Proximal Policy Optimization Algorithms 2017 · Vaswani et al. 2017 · Attention Is All You Need 2018 · Devlin et al. 2018 · BERT: Pre-training of Deep Bidirectional Transformers for Language U… 2018 · Fan et al. 2018 · Hierarchical Neural Story Generation 2018 · Howard et al. 2018 · Universal Language Model Fine-tuning for Text Classification 2018 · Peters et al. 2018 · Deep contextualized word representations 2018 · Radford et al. 2018 · Improving Language Understanding by Generative Pre-Training 2018 · Sutton et al. 2018 · Reinforcement Learning: An Introduction (2nd edition) 2019 · Holtzman et al. 2019 · The Curious Case of Neural Text Degeneration 2019 · Radford et al. 2019 · Language Models are Unsupervised Multitask Learners 2019 · Raffel et al. 2019 · Exploring the Limits of Transfer Learning with a Unified Text-to-Tex… 2020 · Brown et al. 2020 · Language Models are Few-Shot Learners 2020 · Dosovitskiy et al. 2020 · An Image is Worth 16x16 Words: Transformers for Image Recognition at… 2020 · Kaplan et al. 2020 · Scaling Laws for Neural Language Models 2020 · Xiong et al. 2020 · On Layer Normalization in the Transformer Architecture 2021 · Bommasani et al. 2021 · On the Opportunities and Risks of Foundation Models 2021 · Radford et al. 2021 · Learning Transferable Visual Models From Natural Language Supervision 2021 · Wei et al. 2021 · Finetuned Language Models Are Zero-Shot Learners 2022 · Hoffmann et al. 2022 · Training Compute-Optimal Large Language Models 2022 · Min et al. 2022 · Rethinking the Role of Demonstrations: What Makes In-Context Learnin… 2022 · Olsson et al. 2022 · In-context Learning and Induction Heads 2022 · Ouyang et al. 2022 · Training language models to follow instructions with human feedback 2022 · Radford et al. 2022 · Robust Speech Recognition via Large-Scale Weak Supervision 2022 · Wei et al. 2022 · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models 2022 · Wei et al. 2022 · Emergent Abilities of Large Language Models 2023 · Delétang et al. 2023 · Language Modeling Is Compression 2023 · OpenAI et al. 2023 · GPT-4 Technical Report 2023 · Schaeffer et al. 2023 · Are Emergent Abilities of Large Language Models a Mirage? 2023 · Touvron et al. 2023 · LLaMA: Open and Efficient Foundation Language Models 2025 · DeepSeek-AI et al. 2025 · DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforc… 1. What Is Artificial Intelligence? 2. Math Toolkit 3. Machine Learning 4. Neural Networks 5. Vision, Speech & Reinforcement Learning 6. Language Before Transformers 7. Transformers 8. Rise of Large Language Models Beyond (later chapters) 1901 1943 1948 1950 1955 1958 1962 1966 1968 1976 1980 1982 1986 1989 1992 1994 1995 2001 2003 2006 2009 2010 2012 2015 2020 2025 Bengio et al. 2003: A Neural Probabilistic Language Model Mikolov et al. 2013: Efficient Estimation of Word Representations in Vector Space Mikolov et al. 2013: Distributed Representations of Words and Phrases and their Compositionality Pennington et al. 2014: GloVe: Global Vectors for Word Representation Cho et al. 2014: Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation Sutskever et al. 2014: Sequence to Sequence Learning with Neural Networks Sutskever et al. 2014 Bahdanau et al. 2014: Neural Machine Translation by Jointly Learning to Align and Translate Bahdanau et al. 2014 Vaswani et al. 2017: Attention Is All You Need Vaswani et al. 2017 He et al. 2015: Deep Residual Learning for Image Recognition He et al. 2015 Ba et al. 2016: Layer Normalization Ba et al. 2016 Devlin et al. 2018: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding Devlin et al. 2018 Raffel et al. 2019: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer Raffel et al. 2019 Xiong et al. 2020: On Layer Normalization in the Transformer Architecture Xiong et al. 2020 Sennrich et al. 2015: Neural Machine Translation of Rare Words with Subword Units Brown et al. 2020: Language Models are Few-Shot Learners Brown et al. 2020 Kaplan et al. 2020: Scaling Laws for Neural Language Models Kaplan et al. 2020 Hoffmann et al. 2022: Training Compute-Optimal Large Language Models Kingma et al. 2014: Adam: A Method for Stochastic Optimization Shannon 1948: A Mathematical Theory of Communication Robbins et al. 1951: A Stochastic Approximation Method Kullback et al. 1951: On Information and Sufficiency Ruder 2016: An overview of gradient descent optimization algorithms Loshchilov et al. 2017: Decoupled Weight Decay Regularization McCulloch et al. 1943: A logical calculus of the ideas immanent in nervous activity Turing 1950: Computing Machinery and Intelligence McCarthy et al. 1955: A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence, August 31, 1955 Rosenblatt 1958: The perceptron: A probabilistic model for information storage and organization in the brain. Samuel 1959: Some Studies in Machine Learning Using the Game of Checkers Weizenbaum 1966: ELIZA—a computer program for the study of natural language communication between man and machine Hart et al. 1968: A Formal Basis for the Heuristic Determination of Minimum Cost Paths Newell et al. 1976: Computer science as empirical inquiry Rumelhart et al. 1986: Learning representations by back-propagating errors Krizhevsky et al. 2012: ImageNet Classification with Deep Convolutional Neural Networks Mnih et al. 2015: Human-level control through deep reinforcement learning Silver et al. 2016: Mastering the game of Go with deep neural networks and tree search Christiano et al. 2017: Deep reinforcement learning from human preferences Radford et al. 2021: Learning Transferable Visual Models From Natural Language Supervision Radford et al. 2021 Wei et al. 2022: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models Ouyang et al. 2022: Training language models to follow instructions with human feedback DeepSeek-AI et al. 2025: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning Domingos 2012: A few useful things to know about machine learning Kaufman et al. 2012: Leakage in data mining Cortes et al. 1995: Support-vector networks Breiman 2001: Random Forests Tibshirani 1996: Regression Shrinkage and Selection Via the Lasso Pearson 1901: On lines and planes of closest fit to systems of points in space Lloyd 1982: Least squares quantization in PCM Hornik et al. 1989: Multilayer feedforward networks are universal approximators Cybenko 1989: Approximation by superpositions of a sigmoidal function Bengio et al. 1994: Learning long-term dependencies with gradient descent is difficult Hochreiter et al. 1997: Long Short-Term Memory LeCun et al. 1998: Gradient-based learning applied to document recognition Glorot et al. 2010: Understanding the difficulty of training deep feedforward neural networks Srivastava et al. 2014: Dropout: A Simple Way to Prevent Neural Networks from Overfitting Ioffe et al. 2015: Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift He et al. 2015: Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification Hendrycks et al. 2016: Gaussian Error Linear Units (GELUs) Hubel et al. 1962: Receptive fields, binocular interaction and functional architecture in the cat's visual cortex Fukushima 1980: Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position LeCun et al. 1989: Backpropagation Applied to Handwritten Zip Code Recognition Deng et al. 2009: ImageNet: A large-scale hierarchical image database Ren et al. 2015: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks Ronneberger et al. 2015: U-Net: Convolutional Networks for Biomedical Image Segmentation Dosovitskiy et al. 2020: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Dosovitskiy et al. 2020 Graves et al. 2006: Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks Hinton et al. 2012: Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups Oord et al. 2016: WaveNet: A Generative Model for Raw Audio Radford et al. 2022: Robust Speech Recognition via Large-Scale Weak Supervision Radford et al. 2022 Sutton et al. 2018: Reinforcement Learning: An Introduction (2nd edition) Watkins et al. 1992: Q-learning Williams 1992: Simple statistical gradient-following algorithms for connectionist reinforcement learning Mnih et al. 2016: Asynchronous Methods for Deep Reinforcement Learning Schulman et al. 2017: Proximal Policy Optimization Algorithms Peters et al. 2018: Deep contextualized word representations Howard et al. 2018: Universal Language Model Fine-tuning for Text Classification Radford et al. 2018: Improving Language Understanding by Generative Pre-Training Radford et al. 2018 Radford et al. 2019: Language Models are Unsupervised Multitask Learners Fan et al. 2018: Hierarchical Neural Story Generation Holtzman et al. 2019: The Curious Case of Neural Text Degeneration Bommasani et al. 2021: On the Opportunities and Risks of Foundation Models Wei et al. 2021: Finetuned Language Models Are Zero-Shot Learners Min et al. 2022: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? Olsson et al. 2022: In-context Learning and Induction Heads Wei et al. 2022: Emergent Abilities of Large Language Models Schaeffer et al. 2023: Are Emergent Abilities of Large Language Models a Mirage? Delétang et al. 2023: Language Modeling Is Compression Touvron et al. 2023: LLaMA: Open and Efficient Foundation Language Models OpenAI et al. 2023: GPT-4 Technical Report built on (ancestors) built on it (descendants)Larger dots: essential papers. Click a dot, or pick from the list.
2017 · essential
Attention Is All You Need
Introduced the Transformer — the architecture behind BERT, GPT and nearly every modern large language model, and later adapted to vision, audio and more.
Built directly on
2014 · Sutskever et al. 2014 2014 · Bahdanau et al. 2014 2015 · He et al. 2015 2016 · Ba et al. 2016 Directly influenced
2018 · Devlin et al. 2018 2019 · Raffel et al. 2019 2020 · Brown et al. 2020 2020 · Kaplan et al. 2020 2020 · Xiong et al. 2020 2021 · Radford et al. 2021 2020 · Dosovitskiy et al. 2020 2022 · Radford et al. 2022 2018 · Radford et al. 2018 Links record direct influence as described in each paper and in the chapters; they are not a complete citation graph. Browse every paper with its summary in the paper library .