Hands-On LLM Serving and Optimization: Hosting LLMs at Scale
Language: English
Published by O'Reilly Media, 2026
- Softcover
- New

Seller: Rarewaves USA, HEBRON, KY, U.S.A.Rarewaves USA
AbeBooks seller since June 10, 2025
Condition: New
US$ 67.98
Quantity: Over 20 available
Add to basketSeller Inventory # LU-9798341621497
- Title
- Hands-On LLM Serving and Optimization: Hosting LLMs at Scale
- Author
- Chi Wang, Peiheng Hu
- Publisher
- O'Reilly Media
- Publication year
- 2026
- Condition
- New
- Binding
- Paperback
- Language
- English
- ISBN 13
- 9798341621497
- Item weight
- 594 grams
- Dimensions
- 17.78 x 5.08 x 23.34 cm
Large language models (LLMs) are the reasoning engines of modern AI. Today, a major inflection point has arrived: as the world races to deploy AI at scale, model inference has moved to the center of the stack. Welcome to the inference era.
Without proper optimization, however, LLMs can be expensive and slow to serve. Hands-On LLM Serving and Optimization is a comprehensive guide to the complexities of deploying and optimizing LLMs at scale.
In this hands-on, engineering-focused book, authors Chi Wang and Peiheng Hu combine practical examples, code, and strategies for building robust, performant, and cost-efficient AI token factories. Whether you're building the LLM inference infrastructure or the applications that consume it, a deep understanding of LLM serving will make you a more effective, future-ready engineer as AI transforms how we work and build.
- Learn the foundations of model serving with core concepts, design paradigms, and industry best practices
- Understand the common challenges of hosting LLMs at scale
- Balance latency and throughput to meet the demands of AI applications and business requirements
- Host LLMs cost-effectively with practical, code-backed techniques
"Synopsis" may belong to another edition of this title.
About the Author
Peiheng Hu is an accomplished machine learning engineer with over 10 years of industry experience and expertise in building large-scale AI systems. He currently works at NVIDIA, where he focuses on the cutting-edge distributed LLM inference, pushing the boundaries of high-performance inference engines on the latest NVIDIA GPUs. He holds a master of science in computational science and engineering from Harvard University and a bachelor of science in industrial engineering operations research from Georgia Institute of Technology. Previously, Peiheng served as a principal member of technical staff at Salesforce, where he led the development of the company's only unified serving platform, handling thousands of per-tenant models and LLM optimizations for Agentforce that saved millions in AI infrastructure expenses. Prior to that, he was a senior ML engineer at Microsoft Azure, where he architected distributed ML processing solutions for cloud security detection and analytics, handling billions of transactions per hour.
"About the title" may belong to another edition of this title.
Shipping rates within U.S.A.
| Item | 10 to 13 business days | 10 to 13 business days |
|---|---|---|
| First item | US$ 0.00 | US$ 0.00 |
Payment methods
Seller's business information
Rarewaves USA
10100 West Sample Road, Ste 101
Coral Springs, FL U.S.A. 33065
Shipping terms
Please note that we do not offer Priority shipping to any country.
We currently do not ship to the below countries:
Afghanistan
Bhutan
Brazil
Brunei Darussalam
Channel Islands
Chile
Israel
Lao
Mexico
Russian Federation
Saudi Arabia
South Africa
Yemen
Please do not attempt to place orders with any of these countries as a ship to address - they will be cancelled.