Ship fast, private, and efficient AI in the browser with ONNX Runtime Web, WebGPU, and a proven pipeline from training to production.
Developers face real constraints in the browser, from cross origin isolation and CSP to GPU limits, storage quotas, and variable device performance. This book gives you a practical end to end system that handles those constraints while delivering low latency features users can trust.
You will learn when client side inference is the right call, how to export and optimize models, and how to run them reliably with WebGPU and WASM using clean fallbacks. The result is a codebase that is faster to maintain, easier to ship, and ready for production.
This is a code heavy guide, with working examples that show IO binding, fp16 kernels, manifests, service workers, feature probes, and complete project wiring so you can ship real products.
Get the guide that turns browser AI from a demo into a dependable product, grab your copy today.
"synopsis" may belong to another edition of this title.
Seller: GreatBookPrices, Columbia, MD, U.S.A.
Condition: New. Seller Inventory # 51864002-n
Seller: Grand Eagle Retail, Bensenville, IL, U.S.A.
Paperback. Condition: new. Paperback. Ship fast, private, and efficient AI in the browser with ONNX Runtime Web, WebGPU, and a proven pipeline from training to production.Developers face real constraints in the browser, from cross origin isolation and CSP to GPU limits, storage quotas, and variable device performance. This book gives you a practical end to end system that handles those constraints while delivering low latency features users can trust.You will learn when client side inference is the right call, how to export and optimize models, and how to run them reliably with WebGPU and WASM using clean fallbacks. The result is a codebase that is faster to maintain, easier to ship, and ready for production.Decide where inference should run using clear cost, privacy, and latency trade offsExport PyTorch models to ONNX with external data to handle 2 GiB limitsConvert and optimize graphs into ORT format and apply mixed precision fp16 safelyUse ONNX Runtime Web SessionOptions, IO binding, and device tensors to keep data on GPUApply graph capture on WebGPU for static shapes and plan around binding size limitsReach stable CPU performance with WASM threads and SIMD through cross origin isolationProbe WebGPU features including shader f16, subgroups, and timestamp queriesSelect providers with a backend matrix that adapts across WebGPU, WebNN, and WASMBuild tokenizer workflows and pipelines with Transformers.js using device webgpuImplement preprocessing and postprocessing for images and audio with codecs and batchingCache large weights in IndexedDB and OPFS with quota checks and eviction handlingVersion and validate assets using manifests, ETags, and integrity checksStream and lazy load sharded weights with HTTP Range for faster first useHandle large models with KV cache tiling and binding size aware layoutsPartition graphs between WebGPU and WASM for selective fallbacks on weaker devicesBuild real projects end to end, including Stable Diffusion Turbo with fp16 weightsWire Whisper Tiny for streaming capture with VAD for robust speech inputShip real time background removal with camera compositing for the webDeliver CLIP image search with local embeddings and an IndexedDB vector indexSet headers and CSP correctly, avoid COEP credentialless pitfalls, and keep isolationProfile performance with ORT logs and the WebGPU inspector to remove bottlenecksMigrate cleanly from ONNX.js and TensorFlow.js to ORT Web without breaking flowsStand up production telemetry, error boundaries, and alerting that respect privacyPlan costs for CDN egress, caching, and storage, with practical distribution strategiesFuture proof with feature detection for WebNN and NPUs and a maintenance roadmapThis is a code heavy guide, with working examples that show IO binding, fp16 kernels, manifests, service workers, feature probes, and complete project wiring so you can ship real products.Get the guide that turns browser AI from a demo into a dependable product, grab your copy today. This item is printed on demand. Shipping may be from multiple locations in the US or from the UK, depending on stock availability. Seller Inventory # 9798273164123
Seller: California Books, Miami, FL, U.S.A.
Condition: New. Print on Demand. Seller Inventory # I-9798273164123
Seller: GreatBookPrices, Columbia, MD, U.S.A.
Condition: As New. Unread book in perfect condition. Seller Inventory # 51864002
Seller: PBShop.store US, Wood Dale, IL, U.S.A.
PAP. Condition: New. New Book. Shipped from UK. Established seller since 2000. Seller Inventory # L2-9798273164123
Seller: PBShop.store UK, Fairford, GLOS, United Kingdom
PAP. Condition: New. New Book. Shipped from UK. Established seller since 2000. Seller Inventory # L2-9798273164123
Quantity: Over 20 available
Seller: GreatBookPricesUK, Woodford Green, United Kingdom
Condition: New. Seller Inventory # 51864002-n
Quantity: Over 20 available
Seller: GreatBookPricesUK, Woodford Green, United Kingdom
Condition: As New. Unread book in perfect condition. Seller Inventory # 51864002
Quantity: Over 20 available
Seller: CitiRetail, Stevenage, United Kingdom
Paperback. Condition: new. Paperback. Ship fast, private, and efficient AI in the browser with ONNX Runtime Web, WebGPU, and a proven pipeline from training to production.Developers face real constraints in the browser, from cross origin isolation and CSP to GPU limits, storage quotas, and variable device performance. This book gives you a practical end to end system that handles those constraints while delivering low latency features users can trust.You will learn when client side inference is the right call, how to export and optimize models, and how to run them reliably with WebGPU and WASM using clean fallbacks. The result is a codebase that is faster to maintain, easier to ship, and ready for production.Decide where inference should run using clear cost, privacy, and latency trade offsExport PyTorch models to ONNX with external data to handle 2 GiB limitsConvert and optimize graphs into ORT format and apply mixed precision fp16 safelyUse ONNX Runtime Web SessionOptions, IO binding, and device tensors to keep data on GPUApply graph capture on WebGPU for static shapes and plan around binding size limitsReach stable CPU performance with WASM threads and SIMD through cross origin isolationProbe WebGPU features including shader f16, subgroups, and timestamp queriesSelect providers with a backend matrix that adapts across WebGPU, WebNN, and WASMBuild tokenizer workflows and pipelines with Transformers.js using device webgpuImplement preprocessing and postprocessing for images and audio with codecs and batchingCache large weights in IndexedDB and OPFS with quota checks and eviction handlingVersion and validate assets using manifests, ETags, and integrity checksStream and lazy load sharded weights with HTTP Range for faster first useHandle large models with KV cache tiling and binding size aware layoutsPartition graphs between WebGPU and WASM for selective fallbacks on weaker devicesBuild real projects end to end, including Stable Diffusion Turbo with fp16 weightsWire Whisper Tiny for streaming capture with VAD for robust speech inputShip real time background removal with camera compositing for the webDeliver CLIP image search with local embeddings and an IndexedDB vector indexSet headers and CSP correctly, avoid COEP credentialless pitfalls, and keep isolationProfile performance with ORT logs and the WebGPU inspector to remove bottlenecksMigrate cleanly from ONNX.js and TensorFlow.js to ORT Web without breaking flowsStand up production telemetry, error boundaries, and alerting that respect privacyPlan costs for CDN egress, caching, and storage, with practical distribution strategiesFuture proof with feature detection for WebNN and NPUs and a maintenance roadmapThis is a code heavy guide, with working examples that show IO binding, fp16 kernels, manifests, service workers, feature probes, and complete project wiring so you can ship real products.Get the guide that turns browser AI from a demo into a dependable product, grab your copy today. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability. Seller Inventory # 9798273164123
Quantity: 1 available