Items related to NVIDIA GPU Microarchitecture and Optimization: Mastering...

NVIDIA GPU Microarchitecture and Optimization: Mastering CUDA Kernels, Parallel Execution, Memory Hierarchy, Warp Scheduling, Profiling, and Workload Tuning (The Modern Engineer's Library) - Softcover

Book 14 of 14: The Modern Engineer's Library

Frampton, Phyllis H.

 
9798175045759: NVIDIA GPU Microarchitecture and Optimization: Mastering CUDA Kernels, Parallel Execution, Memory Hierarchy, Warp Scheduling, Profiling, and Workload Tuning (The Modern Engineer's Library)

Synopsis

Have you ever pushed a CUDA kernel into production, watched the profiler, and quietly wondered why it's only using a fraction of the GPU you paid for?

You're not alone. Most engineers learn just enough parallel programming to get something working, then hit a wall the moment they need it to actually perform. The gap between "it runs" and "it's fast" isn't a syntax problem. It's a hardware understanding problem. And that's exactly what this book is built to close.

This isn't another surface-level tutorial that shows you how to launch a kernel and calls it a day. This is a complete, ground-up education in how modern GPU hardware actually thinks, so that every optimization decision you make afterward is grounded in real understanding instead of guesswork and copy-pasted tricks from forum posts.

So let me ask you a few honest questions:

Do you actually know what a warp is doing when your threads diverge, or are you just hoping the compiler handles it? Can you look at a kernel's memory access pattern and predict, before you even run it, whether it'll hit the bandwidth ceiling? Do you know the difference between a kernel that's compute bound and one that's memory bound, and why that distinction should completely change how you optimize it? If any of those questions gave you pause, this book was written for you.

Inside, you'll build a real mental model of the hardware itself: streaming multiprocessors, the memory hierarchy from registers down to global memory, and the scheduling logic that determines whether your GPU is actually busy or just spinning its wheels. From there, the book walks you through the full programming model step by step, thread and block organization, warp-level execution, and the exact indexing patterns that separate a kernel that merely works from one that's genuinely efficient.

You'll spend real time inside the memory system, because that's where most performance is won or lost. Coalesced access, shared memory tiling, bank conflicts, register pressure, and the caching behavior that quietly shapes every kernel you write. You'll learn to read a profiler's output the way a mechanic reads an engine, spotting exactly where cycles are being wasted and why.

Then the book moves into territory a lot of guides skip entirely: concurrency and streams, multi-GPU scaling, tensor core programming, mixed precision training, and quantization for inference. If you're working anywhere near machine learning infrastructure, scientific computing, or large-scale data processing, these chapters alone are worth the price of admission.

And because knowing a technique isn't the same as knowing when to use it, the book closes with hands-on capstone projects, a custom convolution kernel, a scientific simulation, a full inference pipeline, and a multi-GPU scaling exercise, each one built the same way real optimization work actually happens: start naive, profile honestly, fix what the data tells you to fix, and measure again.

Here's what you won't find in this book: recycled documentation, filler chapters, or vague advice to "just try different block sizes." Every technique is explained with the reasoning behind it, so that when the next generation of hardware ships, you're not starting over. You'll already know how to think about it.

Who is this actually for? Software engineers moving into performance-critical roles. Machine learning practitioners who want to understand what's happening beneath their framework of choice. Computer science students who want an edge going into systems or infrastructure roles. And working GPU programmers who know the basics but have never had the hardware explained to them properly, the way this book explains it.

If you're ready to stop guessing and start actually understanding the hardware you're programming for, this is where that starts. Scroll up, and let's get into it.

"synopsis" may belong to another edition of this title.