Isbn: 9798192589724 - Engineering Big Data Pipelines with Hadoop and Spark: Store, Process, and Analyze Large-scale Datasets Using Hadoop, Spark, and Pyspark (5 results)

ISBN: 
Refine with Advanced Search

Refine your search

  • Books (5)

  • New (5)

to

Custom price range (US$)

to

  • Language: English

    Published by Amazon Digital Services LLC - Kdp, 2026

    9798192589724

    • Softcover

    Seller: PBShop.store US, Wood Dale, IL, U.S.A.PBShop.store US

    5-star seller
    Contact seller

    Condition: New

    US$ 24.33

     Free Shipping 
    Ships within U.S.A.

    Quantity: Over 20 available

    PAP. Condition: New. New Book. Shipped from UK. Established seller since 2000.

  • Language: English

    Published by Independently published, 2026

    9798192589724

    • Softcover

    Seller: PBShop.store UK, Fairford, GLOS, United KingdomPBShop.store UK

    5-star seller
    Contact seller

    Condition: New

    US$ 22.93

    US$ 4.36 shipping 
    Ships from United Kingdom to U.S.A.

    Quantity: Over 20 available

    PAP. Condition: New. New Book. Shipped from UK. Established seller since 2000.

  • Language: English

    Published by Amazon Digital Services LLC - Kdp Aug 2026, 2026

    9798192589724

    • Softcover

    Seller: AHA-BUCH GmbH, Einbeck, GermanyAHA-BUCH GmbH

    5-star seller
    Contact seller

    Condition: New

    US$ 29.32

    US$ 39.38 shipping 
    Ships from Germany to U.S.A.

    Quantity: 2 available

    Taschenbuch. Condition: Neu. Neuware.

  • Language: English

    Published by Independently published, 2026

    9798192589724

    • Softcover
    • Print on Demand

    Seller: California Books, Miami, FL, U.S.A.California Books

    5-star seller
    Contact seller

    Condition: New

    US$ 24.00

     Free Shipping 
    Ships within U.S.A.

    Quantity: Over 20 available

    Condition: New. Print on Demand.

  • Language: English

    Published by Independently Published, 2026

    9798192589724

    • Softcover
    • Print on Demand

    Seller: CitiRetail, Stevenage, United KingdomCitiRetail

    5-star seller
    Contact seller

    Condition: New

    US$ 27.94

    US$ 48.99 shipping 
    Ships from United Kingdom to U.S.A.

    Quantity: 1 available

    Paperback. Condition: new. Paperback. Learn how to build practical big data pipelines with Hadoop, Apache Spark, and PySpark.As datasets grow beyond the practical limits of a single computer, organizations need reliable ways to store, process, clean, transform, and analyze information at scale. Engineering Big Data Pipelines with Hadoop and Spark provides a practical introduction to the tools and workflows used to solve these problems.This beginner friendly guide takes you from the foundations of big data and distributed computing to complete processing pipelines using Hadoop, Spark, and Python. You will learn not only what the technologies do, but how their individual components fit together in a repeatable data engineering workflow.Inside this book, you will learn how to: Understand big data, distributed computing, batch processing, and streamingUnderstand the Hadoop ecosystem and how its major components work togetherWork with HDFS for distributed storageUnderstand MapReduce and parallel batch processingUse YARN for cluster resource managementSet up Hadoop, Spark, and PySpark on a practical local environmentCreate and work with PySpark applicationsUnderstand RDDs, transformations, actions, and Spark executionWork with Spark DataFrames for large scale data processingQuery large datasets using Spark SQLRead and clean CSV, JSON, and Parquet datasetsTransform large datasets using PySparkUnderstand partitions, shuffling, caching, persistence, and broadcast joinsRead Spark execution plans and use the Spark Web UI for performance analysisProcess continuously arriving data with Structured StreamingBuild introductory machine learning workflows using Spark MLlibDevelop classification and regression models with SparkSave and reload machine learning pipelinesBuild and monitor a complete end to end big data processing projectThe book emphasizes a practical workflow rather than memorizing commands. You will learn to move data through a complete pipeline: store it, inspect it, clean it, transform it, analyze it, optimize the processing job, and save the results.Practical examples use familiar business data including sales transactions, customer orders, product records, website logs, incoming events, and customer behavior data. The final project combines the major skills developed throughout the book by moving raw transaction data into HDFS, processing it with PySpark, analyzing it with Spark SQL, optimizing the job, and saving the results as Parquet.You do not need previous experience with Hadoop, Spark, distributed systems, or data engineering. Basic computer skills are enough to begin, while some familiarity with Python and SQL can be helpful as you progress into PySpark and Spark SQL.Whether you are a student learning data processing, a Python user working with larger datasets, an analyst exploring distributed computing, a software developer moving into data engineering, or a technical professional preparing to work with Hadoop and Spark, this book provides a structured path from fundamentals to practical implementation.Build the skills to design, process, optimize, and understand large scale data pipelines with Hadoop, Spark, and PySpark. This item is printed on demand. Shipping may be from our UK warehouse or from our Australian or US warehouses, depending on stock availability. …