-
Large-Scale LLM Training Systems - How and Why
How large-scale LLM training systems are designed, from the 16 bytes per parameter and mixed precision to collectives, ZeRO and FSDP, tensor, sequence, context, expert, and pipeline parallelism, 4D composition, MFU, and what breaks at scale
-
LLM Inference Systems - From the Roofline to vLLM Internals
How LLM serving engines work, from prefill, decode, and the roofline to paged KV caches, scheduling, speculative decoding, parallelism, and the vLLM V1 internals that implement them
-
LayerNorm and RMSNorm - Forward Pass, Backward Pass, and Every Derivative
Deriving the normalization backward pass from first principles, why it has exactly three terms, and why RMSNorm has two
-
einsum - From Index Notation to Code
The index-form equations you already wrote are einsum strings; the translation is mechanical, and it removes every transpose decision
-
A Simple MLP - Forward Pass, Backward Pass, and Every Derivative
Deriving every gradient in a two-layer MLP from first principles, checking shapes at each step, and turning it into working code