Skip to content
A post-mortem of deep learning, 2012 to 2026

Why the breakthroughs actually worked

Every part of a modern Transformer has a history. Take one apart, trace how each block evolved from AlexNet to today's frontier models, and rerun the experiments that show why each replacement won.

decoder-only LM(B, T, d)
Feed-forward share of params 65% Val loss, no residual 1.74 → 3.35 Each experiment, laptop CPU < 10 min
Level
Introductory to intermediate
You should know
Python and a little linear algebra
Format
Readable chapters with runnable notebooks
Hardware
A laptop CPU, or a free Colab session

The evolution tree

13 lineages · 2010 to 2026

Each row is one branch of a building block's history, placed by year. Hover or tap a step to see what changed and everything that led to it.

Six blocks, six lineages

Part II · 6 lineages

Each chapter follows one block from its first form to today's, with vision and language as two branches that merge at the Transformer.

How the weights get there

Part III · 4 lineages

The loss, the optimizer, the schedule and the starting point: the parts of a model that never appear in its diagram.

Breakthroughs that are not one block

Part IV · 3 lineages

Scale, post-training and inference changed what the same architecture can do.

Objectives

Each chapter opens with what you will be able to do by the end, and closes with questions to check it.

Diagram and equations

Every architecture drawn with its tensor shapes, and the key equations written out next to the code.

Code

30 to 100 lines of PyTorch, laid out to match the diagram block by block.

Predict, then run

Before each experiment you guess the outcome, then a small TinyShakespeare run shows the effect with and without the idea.

Post-mortem

What the paper claimed, what later work showed, and how sure we are: established likely speculative

The tree as plain text

A * marks a major step (a full deep dive); everything else is a short note. Years are approximate until each chapter checks them against the papers.

Modern Transformer (2026)
│
├── ARCHITECTURE
│   ├── Tokens & embeddings
│   │   ├── language: one-hot → word2vec* (2013) → BPE (2016) → SentencePiece, byte-level BPE (2018-19)
│   │   ├── vision:   raw pixels into CNNs (2012) → image patches, ViT* (2020)
│   │   └── merged:   multimodal tokens (2022-24)
│   ├── Position encoding
│   │   ├── implicit: convolution windows (vision), recurrence order (language)
│   │   └── explicit: sinusoidal (2017) → learned (2018) → relative (2018) → RoPE* (2021), ALiBi (2021) → RoPE scaling (2023)
│   ├── Token mixing
│   │   ├── vision:   convolution, AlexNet* (2012) → VGG, Inception (2014)
│   │   ├── language: RNN → LSTM seq2seq* (2014) → Bahdanau attention* (2014)
│   │   ├── merged:   multi-head self-attention, Transformer* (2017) → ViT (2020)
│   │   ├── cheaper:  MQA (2019) → GQA (2023) → MLA (2024); FlashAttention (2022); sliding window (2020-23)
│   │   └── rival:    state-space models, Mamba (2023)
│   ├── Channel mixing (feed-forward)
│   │   ├── activations*: sigmoid, tanh → ReLU (2010-12) → GELU (2016) → Swish (2017) → SwiGLU (2020)
│   │   └── sparsity:     sparsely gated MoE* (2017) → Switch (2021) → Mixtral (2023) → DeepSeekMoE (2024)
│   ├── Normalization
│   │   └── none → BatchNorm* (2015) → LayerNorm (2016) → pre-norm (2019-20) → RMSNorm (2019) → QK-norm (2023)
│   └── Residual connections
│       └── highway nets (2015) → ResNet* (2015) → pre-norm residual stream (2019) → residual-stream view (2021)
│
├── TRAINING RECIPE
│   ├── Objective & loss
│   │   ├── supervised: softmax cross-entropy (AlexNet, 2012)
│   │   ├── language:   skip-gram (2013) → teacher forcing (2014) → masked LM, BERT* / next-token, GPT* (2018)
│   │   ├── vision:     VAE (2013) → GAN* (2014) → diffusion* (2020)
│   │   └── merged:     contrastive image-text, CLIP* (2021)
│   ├── Optimizer
│   │   └── SGD + momentum → AdaGrad (2011), RMSProp (2012) → Adam* (2014) → AdamW (2017) → Lion, Shampoo, Muon (2023-24)
│   ├── Learning-rate schedule
│   │   └── step decay → warmup + cosine* (2016-17) → warmup-stable-decay (2024), schedule-free (2024)
│   └── Initialization & regularization
│       └── Xavier (2010) → He (2015); dropout (2012), augmentation; weight decay; muP (2022)
│
└── BEYOND THE BLOCK
    ├── Scale:           scaling laws* (2020) → Chinchilla (2022); in-context learning* (GPT-3, 2020); emergence (2022)
    ├── Post-training:   instruction tuning (2021-22) → RLHF* (2022) → DPO (2023) → RL on verifiable rewards* (2024-25)
    └── Inference & use: chain of thought (2022) → test-time compute (2024); tool use & agents (2023-26)