| Sep 13, 2026 | DeepSeek's Other Modules - Fine-Grained MoE, Hyper-Connections, Engram, and DSpark, From First Principles |
| Sep 12, 2026 | DeepSeek's Attention and KV Cache - From MLA to CSA2, From First Principles |
| Sep 11, 2026 | Benchmarking DeepSeek-V4.1-Flash on 8x B300 - Where the Step Time Goes |
| Sep 11, 2026 | Serving DeepSeek-V4.1-Flash on vLLM - Architecture, Config, and Where the Time Goes |
| Sep 10, 2026 | How Luminal Works - E-Graphs, Equality Saturation, and Search |
| Sep 02, 2026 | Large-Scale LLM Training Systems - How and Why |
| Aug 23, 2026 | LLM Inference Systems - From the Roofline to vLLM Internals |
| Aug 23, 2026 | LayerNorm and RMSNorm - Forward Pass, Backward Pass, and Every Derivative |
| Aug 23, 2026 | einsum - From Index Notation to Code |
| Aug 20, 2026 | A Simple MLP - Forward Pass, Backward Pass, and Every Derivative |
| Jun 07, 2026 | CuTe DSL fundamentals and primitives [FA3] |
| May 22, 2026 | Graduated from NYU with MSCS |
| May 15, 2026 | FlashAttention 3 - A Worklog[WIP] |
| Apr 18, 2026 | CUTLASS WGMMA on Hopper - Notes |
| Apr 14, 2026 | Investigating Flaky `test_eagle_dp` — Batch Invariance Failure on L4 GPUs |
| Mar 29, 2026 | GEMM Kernel Optimization Notes |
| Mar 25, 2026 | SiLU+Mul+FP8 Block Quant Pattern Matching Pipeline - vLLM Notes |
| Mar 25, 2026 | Fused SiLU+Mul+FP8 Block Quantization CUDA Kernel - vLLM Notes |
| Mar 10, 2026 | Anatomy of a Spark Job Run |
| Feb 13, 2026 | Transformer Block FLOPs & Parameters Calculations |