Browser Machine Learning Mastery (Paperback)
Language: English
Published by Independently Published, 2025
- Softcover
- New

Seller: Grand Eagle Retail, Bensenville, IL, U.S.A.Grand Eagle Retail
AbeBooks seller since October 12, 2005
Condition: New
US$ 33.69
Quantity: 1 available
Add to basketItem description from seller
Paperback. Ship fast, private, and efficient AI in the browser with ONNX Runtime Web, WebGPU, and a proven pipeline from training to production.Developers face real constraints in the browser, from cross origin isolation and CSP to GPU limits, storage quotas, and variable device performance. This book gives you a practical end to end system that handles those constraints while delivering low latency features users can trust.You will learn when client side inference is the right call, how to export and optimize models, and how to run them reliably with WebGPU and WASM using clean fallbacks. The result is a codebase that is faster to maintain, easier to ship, and ready for production.Decide where inference should run using clear cost, privacy, and latency trade offsExport PyTorch models to ONNX with external data to handle 2 GiB limitsConvert and optimize graphs into ORT format and apply mixed precision fp16 safelyUse ONNX Runtime Web SessionOptions, IO binding, and device tensors to keep data on GPUApply graph capture on WebGPU for static shapes and plan around binding size limitsReach stable CPU performance with WASM threads and SIMD through cross origin isolationProbe WebGPU features including shader f16, subgroups, and timestamp queriesSelect providers with a backend matrix that adapts across WebGPU, WebNN, and WASMBuild tokenizer workflows and pipelines with Transformers.js using device webgpuImplement preprocessing and postprocessing for images and audio with codecs and batchingCache large weights in IndexedDB and OPFS with quota checks and eviction handlingVersion and validate assets using manifests, ETags, and integrity checksStream and lazy load sharded weights with HTTP Range for faster first useHandle large models with KV cache tiling and binding size aware layoutsPartition graphs between WebGPU and WASM for selective fallbacks on weaker devicesBuild real projects end to end, including Stable Diffusion Turbo with fp16 weightsWire Whisper Tiny for streaming capture with VAD for robust speech inputShip real time background removal with camera compositing for the webDeliver CLIP image search with local embeddings and an IndexedDB vector indexSet headers and CSP correctly, avoid COEP credentialless pitfalls, and keep isolationProfile performance with ORT logs and the WebGPU inspector to remove bottlenecksMigrate cleanly from ONNX.js and TensorFlow.js to ORT Web without breaking flowsStand up production telemetry, error boundaries, and alerting that respect privacyPlan costs for CDN egress, caching, and storage, with practical distribution strategiesFuture proof with feature detection for WebNN and NPUs and a maintenance roadmapThis is a code heavy guide, with working examples that show IO binding, fp16 kernels, manifests, service workers, feature probes, and complete project wiring so you can ship real products.Get the guide that turns browser AI from a demo into a dependable product, grab your copy today. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability.…
Seller Inventory # 9798273164123
- Title
- Browser Machine Learning Mastery (Paperback)
- Author
- Aura Fenwick
- Publisher
- Independently Published
- Publication year
- 2025
- Condition
- new
- Binding
- Paperback
- Language
- English
- ISBN 13
- 9798273164123
Ship fast, private, and efficient AI in the browser with ONNX Runtime Web, WebGPU, and a proven pipeline from training to production.
Developers face real constraints in the browser, from cross origin isolation and CSP to GPU limits, storage quotas, and variable device performance. This book gives you a practical end to end system that handles those constraints while delivering low latency features users can trust.
You will learn when client side inference is the right call, how to export and optimize models, and how to run them reliably with WebGPU and WASM using clean fallbacks. The result is a codebase that is faster to maintain, easier to ship, and ready for production.
- Decide where inference should run using clear cost, privacy, and latency trade offs
- Export PyTorch models to ONNX with external data to handle 2 GiB limits
- Convert and optimize graphs into ORT format and apply mixed precision fp16 safely
- Use ONNX Runtime Web SessionOptions, IO binding, and device tensors to keep data on GPU
- Apply graph capture on WebGPU for static shapes and plan around binding size limits
- Reach stable CPU performance with WASM threads and SIMD through cross origin isolation
- Probe WebGPU features including shader f16, subgroups, and timestamp queries
- Select providers with a backend matrix that adapts across WebGPU, WebNN, and WASM
- Build tokenizer workflows and pipelines with Transformers.js using device webgpu
- Implement preprocessing and postprocessing for images and audio with codecs and batching
- Cache large weights in IndexedDB and OPFS with quota checks and eviction handling
- Version and validate assets using manifests, ETags, and integrity checks
- Stream and lazy load sharded weights with HTTP Range for faster first use
- Handle large models with KV cache tiling and binding size aware layouts
- Partition graphs between WebGPU and WASM for selective fallbacks on weaker devices
- Build real projects end to end, including Stable Diffusion Turbo with fp16 weights
- Wire Whisper Tiny for streaming capture with VAD for robust speech input
- Ship real time background removal with camera compositing for the web
- Deliver CLIP image search with local embeddings and an IndexedDB vector index
- Set headers and CSP correctly, avoid COEP credentialless pitfalls, and keep isolation
- Profile performance with ORT logs and the WebGPU inspector to remove bottlenecks
- Migrate cleanly from ONNX.js and TensorFlow.js to ORT Web without breaking flows
- Stand up production telemetry, error boundaries, and alerting that respect privacy
- Plan costs for CDN egress, caching, and storage, with practical distribution strategies
- Future proof with feature detection for WebNN and NPUs and a maintenance roadmap
This is a code heavy guide, with working examples that show IO binding, fp16 kernels, manifests, service workers, feature probes, and complete project wiring so you can ship real products.
Get the guide that turns browser AI from a demo into a dependable product, grab your copy today.
"Synopsis" may belong to another edition of this title.
Grand Eagle Retail
Bensenville, IL, U.S.A.
AbeBooks seller since October 12, 2005
Shipping rates within U.S.A.
| Item | 6 to 14 business days | 6 to 16 business days |
|---|---|---|
| First item | US$ 0.00 | US$ 0.00 |
Payment methods
Seller's business information
APOLLO ONLINE CORP.
605 Geddes Street
Wilmington, DE U.S.A. 19805
Terms of sale
We guarantee the condition of every book as it¿s described on the Abebooks web sites. If you¿ve changed
your mind about a book that you¿ve ordered, please use the Ask bookseller a question link to contact us
and we¿ll respond within 2 business days.
Books ship from California and Michigan.
Shipping terms
Orders usually ship within 2 business days. All books within the US ship free of charge. Delivery is 4-14 business days anywhere in the United States.
Books ship from California and Michigan.
If your book order is heavy or oversized, we may contact you to let you know extra shipping is required.