- Ternary Weights, 3-Bit KV Caches, and the Limits of Quantization | Bangalore Paper Club - YouTube
- Machine Learning Engineering Open Book github.com/stas00
- Machine Learning: LLM/VLM Training and Engineering by Stas Bekman stasosphere.com
- Understanding LSTM Networks colah.github.io
- Autograd - automatically differentiate native Python and Numpy code github.com/HIPS/autograd
- IOP Systems blogs iop.systems
- E2E pipeline to visualise how LLMs process prompts github.com/taylorsatula
- Attention Is All You Need arxiv.org
- Neural Machine Translation by Jointly Learning to Align and Translate arxiv.org research.google.com
- Learning representations by back-propagating errors www.nature.com
- Efficient Estimation of Word Representations in Vector Space arxiv.orgresearch.google.com
- Scalable Private Search with Wally machinelearning.apple.com
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding research.google.com
- How LLMs Actually Work www.0xkato.xyz
- Can gzip be a language model nathan.rs
- Language Modeling Without Neural Networks nathan.rs
- Tiny LLM - LLM Serving in a Week skyzh.github.io
- An Agent in 100 Lines of Lisp, or How my Prof was Right - Just 25 Years Early thebeach.dev
- Modern GPU Programming for MLSys mlc.ai