Key Papers
The landmark papers behind the milestones — from McCulloch & Pitts in 1943 to the transformer and beyond, each with a link to the original and what it changed.
Grouped by era, and roughly parallel to the history timeline — that page tracks what happened, this one tracks what was written down.
1950s-1960s
- “Computing Machinery and Intelligence” (Turing, 1950) - paper
- Introduced the Turing Test and fundamental questions about machine intelligence
- First serious discussion of whether machines can think
- “A Logical Calculus of Ideas Immanent in Nervous Activity” (McCulloch & Pitts, 1943) - paper
- Laid groundwork for neural networks and computational theory of mind
- Showed how simple neural networks could compute logical functions
1970s-1980s
- “A Framework for Representing Knowledge” (Minsky, 1974) - paper
- Introduced frames as a way to represent knowledge
- Influenced modern knowledge representation systems
- “Learning representations by back-propagating errors” (Rumelhart, Hinton & Williams, 1986) - paper
- Popularized backpropagation for training neural networks
- Enabled practical training of deep neural networks
1990s-2000s
- “Long Short-Term Memory” (Hochreiter & Schmidhuber, 1997) - paper
- Introduced LSTM networks
- Solved the vanishing gradient problem for recurrent neural networks
- “Gradient-Based Learning Applied to Document Recognition” (LeCun et al., 1998) - paper
- Introduced ConvNets and the LeNet architecture
- Pioneered deep learning for computer vision
2010s
- “ImageNet Classification with Deep Convolutional Neural Networks” (Krizhevsky et al., 2012) - paper
- AlexNet paper that sparked the deep learning revolution
- Demonstrated the power of deep CNNs trained on large datasets
- “Attention Is All You Need” (Vaswani et al., 2017) - paper
- Introduced the Transformer architecture
- Foundation for modern language models like GPT and BERT
2018-2020
- “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding” (Devlin et al., 2018) - paper
- Introduced bidirectional context for language understanding
- Set new standards for NLP tasks
- “Language Models are Few-Shot Learners” (Brown et al., 2020) - GPT-3 paper - paper
- Demonstrated emergent abilities in large language models
- Showed scaling effects on model capabilities
- Introduced few-shot learning through prompting
- “High-Resolution Image Synthesis with Latent Diffusion Models” (Rombach et al., 2022) - Stable Diffusion paper - paper
- Made efficient image generation possible on consumer hardware
- Advanced the field of text-to-image generation
2021-2023
- “Constitutional AI: Harmlessness from AI Feedback” (Yuntao et al., 2022) - paper
- Introduced methods for aligning AI systems with human values
- Proposed frameworks for AI safety
- “PaLM: Scaling Language Modeling with Pathways” (Chowdhery et al., 2022) - paper
- Advanced scaling laws for language models
- Demonstrated breakthrough capabilities in reasoning and code generation
- “Training language models to follow instructions with human feedback” (Long et al., 2022) - InstructGPT paper - paper
- Pioneered instruction tuning using human feedback (RLHF)
- Improved model alignment with human intent
- “LLaMA: Open and Efficient Foundation Language Models” (Touvron et al., 2023) - paper
- Demonstrated efficient training of powerful open-source language models
- Sparked widespread development of open LLMs
2018-2020
- “Denoising Diffusion Probabilistic Models” (Ho, Jain & Abbeel, 2020) - paper
- Cast image generation as learning to reverse a fixed Gaussian noising process
- Introduced the simplified epsilon-prediction training objective that made diffusion practical
- “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale” (Dosovitskiy et al., 2020) - ViT paper - paper
- Showed a plain transformer over image patches could match convolutional networks at scale
- Became the vision encoder every multimodal model now uses
2021-2023
- “Score-Based Generative Modeling through Stochastic Differential Equations” (Song et al., 2021) - Score-SDE paper - paper
- Unified score matching and diffusion as discretisations of one SDE
- Introduced the probability-flow ODE and the variance-exploding/preserving view
- “Learning Transferable Visual Models From Natural Language Supervision” (Radford et al., 2021) - CLIP paper - paper
- Aligned images and captions in one embedding space with a contrastive objective
- Enabled zero-shot classification and became the default text encoder for image generation
- “High-Resolution Image Synthesis with Latent Diffusion Models” (Rombach et al., 2022) - Latent diffusion / Stable Diffusion paper - paper
- Moved diffusion into a compressed latent space, cutting cost by an order of magnitude
- Added cross-attention conditioning that made text-to-image controllable
- “Classifier-Free Diffusion Guidance” (Ho & Salimans, 2022) - paper
- Showed guidance could be obtained by interpolating conditional and unconditional predictions
- Removed the need for a separate classifier and became the standard sampling trick
- “Scalable Diffusion Models with Transformers” (Peebles & Xie, 2022) - DiT paper - paper
- Replaced the U-Net with a plain transformer over latent patches
- Introduced adaLN-Zero conditioning and inherited language-model scaling laws
- “Flow Matching for Generative Modeling” (Lipman et al., 2022) - paper
- Trained a velocity field along a chosen probability path with a simple regression loss
- Generalised diffusion and became the objective behind the 2026 image frontier
- “Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow” (Liu, Gong & Liu, 2022) - paper
- Straightened the flow’s integral curves so few-step sampling becomes possible
- Underlies the reflow and shortcut families of fast samplers
- “Robust Speech Recognition via Large-Scale Weak Supervision” (Radford et al., 2022) - Whisper paper - paper
- Trained a single attention encoder-decoder on 680,000 hours of weakly labelled audio
- Set the robustness standard for speech recognition across accents and noise
- “High Fidelity Neural Audio Compression” (Défossez et al., 2022) - EnCodec paper - paper
- Turned audio into a few hundred discrete tokens per second with a residual vector quantiser
- Became the tokenizer that makes audio language modelling possible
- “Sigmoid Loss for Language Image Pre-Training” (Zhai et al., 2023) - SigLIP paper - paper
- Replaced the softmax contrastive loss with a pairwise sigmoid over each image-text pair
- Dropped the global all-gather, making contrastive pretraining cheaper and more stable
- “Flamingo: a Visual Language Model for Few-Shot Learning” (Alayrac et al., 2022) - paper
- Fused a frozen vision encoder into a frozen language model with gated cross-attention
- Handled interleaved images and text without retraining either tower
- “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models” (Li et al., 2023) - paper
- Introduced the Q-Former to query a frozen vision encoder into a fixed set of tokens
- Showed a small trainable bridge could join two frozen models
- “Visual Instruction Tuning” (Liu et al., 2023) - LLaVA paper - paper
- Used a simple linear projector and instruction data to build an open visual assistant
- Established the adapter-and-tune recipe most open VLMs still follow
- “Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers” (Wang et al., 2023) - VALL-E paper - paper
- Treated speech as discrete codec tokens and generated them with a language model
- Achieved zero-shot voice cloning from a three-second prompt
- “Segment Anything” (Kirillov et al., 2023) - SAM paper - paper
- Made segmentation promptable with points, boxes and masks over a 1B-mask dataset
- Turned a dense prediction task into a promptable, groundable interface
- “RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control” (Brohan et al., 2023) - paper
- Discretised robot actions as text tokens on top of a pretrained vision-language model
- Showed web-scale visual knowledge transfers to manipulation
2024-2026
- “Scaling Rectified Flow Transformers for High-Resolution Image Synthesis” (Esser et al., 2024) - Stable Diffusion 3 paper - paper
- Made rectified flow the training objective of a frontier text-to-image model
- Introduced the MMDiT architecture with separate text and image streams fused at attention
- “π₀: A Vision-Language-Action Flow Model for General Robot Control” (Black et al., 2024) - paper
- Added a flow-matching action expert on top of a pretrained vision-language model
- Generated action chunks at high frequency across many robot embodiments