LLM Inference Engineering: From KV Cache to Production Serving: A Practical Guide, Basics to Advanced
Language: English
Published by Independently published, 2026
- Softcover
- New

Seller: California Books, Miami, FL, U.S.A.California Books
AbeBooks seller since October 27, 2023
Condition: New
US$ 25.00
Quantity: Over 20 available
Add to basketItem description from seller
Print on Demand.
Seller Inventory # I-9798178522462
- Title
- LLM Inference Engineering: From KV Cache to Production Serving: A Practical Guide, Basics to Advanced
- Author
- Nainwal, Shashank
- Publisher
- Independently published
- Publication year
- 2026
- Condition
- New
- Binding
- Soft cover
- Language
- English
- ISBN 13
- 9798178522462
Training a model happens once. Serving it happens every time someone uses your product, and that is where most AI budgets go. This book explains, step by step, what really happens when an LLM generates text and how engineers make it fast and affordable at scale.
Starting from a single question, "what happens when a model produces one token?", you will build a complete mental model of modern inference, from first principles to production deployment.
What you will learn
- Why decoding is memory-bound, and how to predict speed and cost with simple napkin math
- The KV cache, PagedAttention, FlashAttention, and GQA/MQA/MLA explained in plain language
- Continuous batching, chunked prefill, speculative decoding, and prompt caching
- Quantization (FP8, INT4, GPTQ, AWQ), distillation, Mixture of Experts, and model routing
- Tensor, pipeline, and expert parallelism, plus disaggregated serving and long context
- How modern serving engines and the GGUF format work, and when to choose each
- How GPUs and TPUs work, and how to compare accelerators for inference
- Benchmarking, observability, cost modeling, and capacity planning
Inside the book
- 27 chapters across 8 parts, from basics to advanced
- 29 original diagrams
- Key takeaways and quiz questions in every chapter, with an answer key
- 3 hands-on labs you can run on a laptop or a single GPU
- 25 interview questions with model answers
- A formula cheat sheet, glossary, and curated list of foundational papers
Who this book is for
Software and ML engineers, platform and MLOps engineers, solution architects, technical product managers, and anyone preparing for AI infrastructure interviews. You need basic Python and a rough idea of what a neural network is. No CUDA or GPU required.
Written by an AI architect and former Amazon Web Services engineer who has built LLM applications serving millions of customers.
"Synopsis" may belong to another edition of this title.
California Books
Miami, FL, U.S.A.
AbeBooks seller since October 27, 2023
Shipping rates within U.S.A.
| Item | 3 to 7 business days | 2 to 5 business days |
|---|---|---|
| First item | US$ 0.00 | US$ 12.00 |
Payment methods
Store description
We have 20 years experience selling books worldwide! Friendly customer support. Your satisfaction guaranteed!
Specialty
All authorized categoriesSeller's business information
Miramar International Services LLC
FL, U.S.A.
Terms of sale
www.californiabooks.com
Shipping terms
www.californiabooks.com