PySpark Cookbook : Over 60 Recipes for Implementing Big Data Processing and Analytics Using Apache Spark and Python
Language: English
Published by Packt Publishing, Limited, 2018
- Softcover
- Used

Seller: Better World Books: West, Reno, NV, U.S.A.Better World Books: West
AbeBooks seller since March 14, 2016
Condition: Used - Good
US$ 17.36
Quantity: 2 available
Add to basketItem description from seller
Pages intact with minimal writing/highlighting. The binding may be loose and creased. Dust jackets/supplements are not included. Stock photo provided. Product includes identifying sticker. Better World Books: Buy Books. Do Good.
Seller Inventory # 49994245-75
- Title
- PySpark Cookbook : Over 60 Recipes for Implementing Big Data Processing and Analytics Using Apache Spark and Python
- Author
- Lee, Denny, Drabas, Tomasz
- Publisher
- Packt Publishing, Limited
- Publication year
- 2018
- Condition
- Good
- Binding
- Soft cover
- Language
- English
- ISBN 10
- 1788835360
- ISBN 13
- 9781788835367
- Item weight
- 1.25 pounds
- Dimensions
- N/A
Combine the power of Apache Spark and Python to build effective big data applications
Key Features
- Perform effective data processing, machine learning, and analytics using PySpark
- Overcome challenges in developing and deploying Spark solutions using Python
- Explore recipes for efficiently combining Python and Apache Spark to process data
Book Description
Apache Spark is an open source framework for efficient cluster computing with a strong interface for data parallelism and fault tolerance. The PySpark Cookbook presents effective and time-saving recipes for leveraging the power of Python and putting it to use in the Spark ecosystem.
You'll start by learning the Apache Spark architecture and how to set up a Python environment for Spark. You'll then get familiar with the modules available in PySpark and start using them effortlessly. In addition to this, you'll discover how to abstract data with RDDs and DataFrames, and understand the streaming capabilities of PySpark. You'll then move on to using ML and MLlib in order to solve any problems related to the machine learning capabilities of PySpark and use GraphFrames to solve graph-processing problems. Finally, you will explore how to deploy your applications to the cloud using the spark-submit command.
By the end of this book, you will be able to use the Python API for Apache Spark to solve any problems associated with building data-intensive applications.
What you will learn
- Configure a local instance of PySpark in a virtual environment
- Install and configure Jupyter in local and multi-node environments
- Create DataFrames from JSON and a dictionary using pyspark.sql
- Explore regression and clustering models available in the ML module
- Use DataFrames to transform data used for modeling
- Connect to PubNub and perform aggregations on streams
Who This Book Is For
The PySpark Cookbook is for you if you are a Python developer looking for hands-on recipes for using the Apache Spark 2.x ecosystem in the best possible way. A thorough understanding of Python (and some familiarity with Spark) will help you get the best out of the book.
Table of Contents
- Spark installation and configuration
- Abstracting data with RDDs
- Abstracting data with DataFrames
- Preparing data for modeling
- Machine Learning with MLLib
- Machine Learning with ML module
- Structured streaming with PySpark
- GraphFrames - Graph Theory with PySpark
"Synopsis" may belong to another edition of this title.
About the Author
Denny Lee is a technology evangelist at Databricks. He is a hands-on data science engineer with 15+ years of experience. His key focuses are solving complex large-scale data problems providing not only architectural direction but hands-on implementation of such systems. He has extensive experience of building greenfield teams as well as being a turnaround/change catalyst. Prior to joining Databricks, he was a senior director of data science engineering at Concur and was part of the incubation team that built Hadoop on Windows and Azure (currently known as HDInsight).
Tomasz Drabas is a data scientist specializing in data mining, deep learning, machine learning, choice modeling, natural language processing, and operations research. He is the author of Learning PySpark and Practical Data Analysis Cookbook. He has a PhD from University of New South Wales, School of Aviation. His research areas are machine learning and choice modeling for airline revenue management.
"About the title" may belong to another edition of this title.
Better World Books: West
Reno, NV, U.S.A.
AbeBooks seller since March 14, 2016
Shipping rates within U.S.A.
| Item | 4 to 8 business days | 3 to 5 business days |
|---|---|---|
| First item | US$ 0.00 | US$ 13.00 |
Payment methods
Store description
Better World Books is a for-profit, socially conscious business and a global online bookseller that collects and sells new and used books online, matching each purchase with a book donation. Each sale generates funds for literacy and education initiatives in the U.S., the UK, and around the world. Since its launch in 2003, Better World Books has raised over $35 million for libraries and literacy, donated over 38 million books, and reused or recycled more than 475 million books.
Seller's business information
Better World Books Marketplace, Inc.
55742 Currant Road
Mishawaka, IN U.S.A. 46545
Terms of sale
Better World Books (BWB) values your satisfaction and offers you returns within thirty (30) days after the estimated delivery date on most items. All returned items must be in the original condition; used items should include the SKU sticker located on the spine or back of the product.
If you have an incomplete, incorrect, or damaged shipment, please contact our Customer Care team via Abebooks contact seller options before proceeding with the return.Please keep in mind that because we deal mostly in used books, any extra components, such as CDs, DVDs, figurines, or access codes are not included.
Shipping terms
Please allow 1-2 business days for order fulfillment.