Inventing LLMs
Chapters
First Steps
Text & Tokenization
N-grams & the Curse
Similarity & Embeddings
Function Approximation
Neural Networks
MLE & Gradient Descent
SGD & Backprop
Regularization
Milestone: Neural LM
Batch Norm
Adam & AdamW
BPE
Residual Connections
Milestone: Neural LM 2.0
Attention
Transformer Architecture
RoPE
RMS Norm
Milestone: GPT
SwiGLU
Mixtures of Experts
Milestone: MoE-GPT
Supervised Fine-Tuning (SFT)
Reinforcement Learning (RL)
Rewards & Reward Models
Policy Gradients
Advantage Estimation
PPO & KL Penalty
Milestone: RL Framework
Book banner image

From scratch

We start from absolute zero and invent every key element of LLM architecture + training from scratch. Intuition for high-school-level math beneficial (general algebra, derivatives, vectors+matrices). Nothing else required.

One coherent story

Simple approach → problem → analysis → less simple approach. Rinse and repeat. Urgent issues first, others tracked for downstream improvements.

Heavy on intuition

First-principle arguments, toy problems, information-oriented reasoning and specific examples wherever appropriate. No top-down explanations, no regurgitation of deprecated framings.

Visual and effective

Fancy visuals and interactive widgets where they actually aid understanding. Highlight boxes for clarity. Text for the main narrative.

Deliberately conceptual

Concepts, concepts, concepts. This is intended as a world-class on-ramp to understanding LLMs for people who want to dig deep. It is not a coding guide.

Start Inventing