[Notes] Designing ML Systems
Gentle Intro to CUDA
[Notes] The Smol Training Playbook
LLM Decoder Architecture Explained
Optimizing Retrieval Augmented Generation
vLLM Server with AWS EKS
vLLM Serve Optimizations
[Notes] LLM Engineer's Handbook
Cells Unpacked
Questions from a Stanford HAI Discussion