Engineering Big Data Pipelines with Hadoop and Spark
Westwood, Nathan
Sold by PBShop.store US, Wood Dale, IL, U.S.A.
AbeBooks Seller since April 7, 2005
New - Soft cover
Condition: New
Ships within U.S.A.
Quantity: Over 20 available
Add to basketSold by PBShop.store US, Wood Dale, IL, U.S.A.
AbeBooks Seller since April 7, 2005
Condition: New
Quantity: Over 20 available
Add to basketNew Book. Shipped from UK. Established seller since 2000.
Seller Inventory # L2-9798192589724
Learn how to build practical big data pipelines with Hadoop, Apache Spark, and PySpark.
As datasets grow beyond the practical limits of a single computer, organizations need reliable ways to store, process, clean, transform, and analyze information at scale. Engineering Big Data Pipelines with Hadoop and Spark provides a practical introduction to the tools and workflows used to solve these problems.
This beginner friendly guide takes you from the foundations of big data and distributed computing to complete processing pipelines using Hadoop, Spark, and Python. You will learn not only what the technologies do, but how their individual components fit together in a repeatable data engineering workflow.
Inside this book, you will learn how to:
The book emphasizes a practical workflow rather than memorizing commands. You will learn to move data through a complete pipeline: store it, inspect it, clean it, transform it, analyze it, optimize the processing job, and save the results.
Practical examples use familiar business data including sales transactions, customer orders, product records, website logs, incoming events, and customer behavior data. The final project combines the major skills developed throughout the book by moving raw transaction data into HDFS, processing it with PySpark, analyzing it with Spark SQL, optimizing the job, and saving the results as Parquet.
You do not need previous experience with Hadoop, Spark, distributed systems, or data engineering. Basic computer skills are enough to begin, while some familiarity with Python and SQL can be helpful as you progress into PySpark and Spark SQL.
Whether you are a student learning data processing, a Python user working with larger datasets, an analyst exploring distributed computing, a software developer moving into data engineering, or a technical professional preparing to work with Hadoop and Spark, this book provides a structured path from fundamentals to practical implementation.
Build the skills to design, process, optimize, and understand large scale data pipelines with Hadoop, Spark, and PySpark.
"About this title" may belong to another edition of this title.
Returns Policy
We ask all customers to contact us for authorisation should they wish to return their order. Orders returned without authorisation may not be credited.
If you wish to return, please contact us within 14 days of receiving your order to obtain authorisation.
Returns requested beyond this time will not be authorised.
Our team will provide full instructions on how to return your order and once received our returns department will process your refund.
Please note the cost to return any...
Books are shipped from UK warehouse. Delivery thereafter is between 4 and 14 business days dependant upon your location - please do contact us with any queries you may have.
| Order quantity | 7 to 14 business days | 7 to 14 business days |
|---|---|---|
| First item | US$ 0.00 | US$ 0.00 |
Delivery times are set by sellers and vary by carrier and location. Orders passing through Customs may face delays and buyers are responsible for any associated duties or fees. Sellers may contact you regarding additional charges to cover any increased costs to ship your items.