- LLM Inference Handbook
- Why Inference is hard.. - YouTube
- SNIA SDC 2025 - KV-Cache Storage Offloading for Efficient Inference in LLMs - YouTube
- LLM Inference Optimisation jonasgeiping.github.io
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference. github.com/wafer-ai
- tiny-vllm: Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM github.com/jmaczan
- RL is Everything, Everywhere, All at Once skypilot.ai
- The Illustrated Transformer jalammar.github.io
- [2605.22850] ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse
- Husky: up to 4.5× faster than MLX
- What are KV caches really?
- Single node inferencing mechanisms
- Inside vLLM: Anatomy of a High-Throughput LLM Inference System | vLLM Blog
- [2309.06180] Efficient Memory Management for Large Language Model Serving with PagedAttention and the vLLM paper
- Orca: A Distributed Serving System for Transformer-Based Generative Models | USENIX
- [2403.02310] Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
- Multi node disaggregation
- Splitwise: Efficient generative LLM inference using phase splitting - Microsoft Research
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
- Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot | USENIX
- [2504.03648] AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure - distributed KV Cache pool
- IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference | USENIX
- Distributed inference