Just for you
Your progress
A plain record of where you are. Mark concepts as you go; revisit anything flagged. Sign in to keep it in your account on every device; without an account it lives only in this browser.
Concepts understood
0 / 109
Chapters completed
0 / 8
Papers read
0 / 88
Videos watched
0 / 43
Experiments done
0 / 25
Must-know concepts understood: 0 of 89. The site contains only published material so far; totals grow as chapters are added. Am I ready to read AI papers? →
Your account
Concepts by status
Not started109
- Activation Functions
- Actor–Critic Methods
- AI Winters
- From MLPs to Transformers: The Architecture Story
- The Artificial Neuron
- Attention
- Audio and Spectrograms
- Autoregressive Next-Token Prediction
- Backpropagation
- From Base Model to ChatGPT
- Batch Normalization
- Conditional Probability and Bayes' Theorem
- Causal Masking
- Probability of Sequences
- The Chain Rule
- Convolutional Neural Networks
- Computational Graphs and Autodiff
- Convolution
- Cross-Entropy Loss
- Data Leakage
- Decision Trees and Random Forests
- Decoding: Greedy, Temperature, Top-k, Top-p
- Deep RL, from Atari to AlphaGo to RLHF
- Derivatives and Gradients
- Detection and Segmentation
- Distribution Shift
- Dot Product
- Dropout
- Embeddings
- Emergent Abilities and the Debate
- Entropy
- Evaluation Metrics for Classifiers
- Expected Value and Variance
- Expert Systems
- Exploration vs Exploitation
- Hand-Crafted Features vs Learned Features
- Features, Labels and Tasks
- Feed-Forward Sublayer (MLP)
- The Fixed-Vector Bottleneck
- The Forward Pass
- Generalization, Overfitting and Underfitting
- GloVe
- GPT-1 → GPT-2 → GPT-3
- Gradient Descent
- ImageNet, AlexNet and ResNet
- Images as Tensors
- In-Context Learning
- Weight Initialization
- k-Means Clustering
- KL Divergence
- The Knowledge-Acquisition Bottleneck
- Knowledge Representation
- Language Modeling
- Layer Normalization
- Supervised, Unsupervised and Self-Supervised Learning
- Linear Regression
- Logic and Rules
- Logistic Regression
- Loss Functions
- LSTMs and GRUs
- MDPs, Policies and Value
- Matrix Multiplication
- Parameters, Tokens and Context Windows
- Momentum and Adam
- Multi-Head Attention
- Multilayer Perceptron (MLP)
- N-Gram Models
- Naive Bayes
- Neural Language Model
- Neural Machine Translation
- One-Hot Encoding
- Principal Component Analysis (PCA)
- The Perceptron
- Perplexity
- Planning
- Policy Gradients
- Pooling and Downsampling
- Positional Encoding
- Why Next-Token Prediction Goes So Far
- Pretrain, Then Fine-Tune
- Pretraining at Scale
- Probability and Distributions
- Q-Learning
- Regularization
- Reinforcement Learning
- Representation Learning
- Residual Connections
- Recurrent Neural Networks
- From Rules to Learning
- Sampling and Uncertainty
- Scaling Laws (Preview)
- Search
- Self-Attention
- Sequence-to-Sequence Models
- Softmax
- Speech Recognition and Synthesis
- Stochastic Gradient Descent (SGD)
- Support Vector Machines
- Symbolic AI
- Tensors and Shapes
- Text as Data
- Tokenization
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- The Turing Test
- Vanishing and Exploding Gradients
- Vectors
- Vision Transformer (ViT)
- Word2Vec
Learning0
Understood0
Revisit0
Your data
Progress lives in this browser's local storage. Nothing is sent anywhere unless you sign in to sync.