Cleaning Data for Effective Data Science: Doing the other 80% of the work with Python, R, and command-line tools
Language: English
Published by Packt Publishing, 2021
- Softcover
- Used

Seller: Goodwill Books, Hillsboro, OR, U.S.A.Goodwill Books
AbeBooks seller since August 25, 2000
Condition: Used - Fair
US$ 17.25
Quantity: 1 available
Add to basketItem description from seller
Fairly worn, but readable and intact. If applicable: Dust jacket, disc or access code may not be included.
Seller Inventory # GICWV.1801071292.A
- Title
- Cleaning Data for Effective Data Science: Doing the other 80% of the work with Python, R, and command-line tools
- Author
- Mertz; David
- Publisher
- Packt Publishing
- Publication year
- 2021
- Condition
- acceptable
- Binding
- Soft cover
- Language
- English
- ISBN 10
- 1801071292
- ISBN 13
- 9781801071291
A comprehensive guide for data scientists to master effective data cleaning tools and techniques
Key Features
- Think about your data intelligently and ask the right questions
- Master data cleaning techniques using hands-on examples belonging to diverse domains
- Work with detailed, commented, well-tested code samples in Python and R
Book Description
In data science, data analysis, or machine learning, most of the effort needed to achieve your actual purpose lies in cleaning your data. Using Python, R, and command-line tools, you will learn the essential cleaning steps performed in every production data science or data analysis pipeline. This book not only teaches you data preparation but also what questions you should ask of your data.
The book dives into the practical application of tools and techniques needed for data ingestion, anomaly detection, value imputation, and feature engineering. It also offers long-form exercises at the end of each chapter to practice the skills acquired.
You will begin by looking at data ingestion of a range of data formats. Moving on, you will impute missing values, detect unreliable data and statistical anomalies, and generate synthetic features that are necessary for successful data analysis and visualization goals.
By the end of this book, you will have acquired a firm understanding of the data cleaning process necessary to perform real-world data science and machine learning tasks.
What you will learn
- Ingest and work with common tabular, hierarchical, and other data formats
- Apply useful rules and heuristics for assessing data quality and detecting bias
- Identify and handle unreliable data and outliers in their many forms
- Impute sensible values into missing data and use sampling to fix imbalances
- Generate synthetic features that help to draw out patterns in your data
- Prepare data competently and correctly for analytic and machine learning tasks
Who this book is for
This book is designed to benefit software developers, data scientists, aspiring data scientists, and students who are interested in data analysis or scientific computing.
Basic familiarity with statistics, general concepts in machine learning, knowledge of a programming language (Python or R), and some exposure to data science are helpful.
The text will also be helpful to intermediate and advanced data scientists who want to improve their rigor in data hygiene and wish for a refresher on data preparation issues.
Table of Contents
- Data Ingestion – Tabular Formats
- Data Ingestion - Hierarchical Formats
- Data Ingestion - Repurposing Data Sources
- The Vicissitudes of Error - Anomaly Detection
- The Vicissitudes of Error - Data Quality
- Rectification and Creation - Value Imputation
- Rectification and Creation - Feature Engineering
- Ancillary Matters - Closure/Glossary
"Synopsis" may belong to another edition of this title.
About the Author
David Mertz, Ph.D. is the founder of KDM Training, a partnership dedicated to educating developers and data scientists in machine learning and scientific computing. He created a data science training program for Anaconda Inc. and was a senior trainer for them. With the advent of deep neural networks, he has turned to training our robot overlords as well.
He previously worked for 8 years with D. E. Shaw Research and was also a Director of the Python Software Foundation for 6 years. David remains co-chair of its Trademarks Committee and Scientific Python Working Group. His columns, Charming Python and XML Matters, were once the most widely read articles in the Python world.
"About the title" may belong to another edition of this title.
Goodwill Books
Hillsboro, OR, U.S.A.
AbeBooks seller since August 25, 2000
Shipping rates within U.S.A.
| Item | 5 to 14 business days | 3 to 6 business days |
|---|---|---|
| First item | US$ 3.99 | US$ 7.00 |
Payment methods
Store description
We are a division of Goodwill Industries of the Columbia Willamette, headquartered in Portland, Oregon, U.S.A. When you buy from Goodwill Books not only will you feel happy about the money you save, you'll feel good knowing your purchase helps Goodwill provide vocational opportunities to people with barriers to employment. For more information please visit our website: www.goodwillbooks.com
Specialty
Fiction, Americana/History, Religion/Spirituality, Children's books, Text books, Art/Humanities, Signed Books, Self-Help, Pacific NorthwestSeller's business information
Goodwill Books
2920 SE Century Blvd
Hillsboro, OR U.S.A. 97123
Terms of sale
All books subject to prior sale. Books shipped via USPS Media Mail or Priority Mail. Extra postage required for heavy books and sets unless otherwise noted. Payment in U.S. dollars via VISA/MASTERCARD/DISCOVER. Sorry, we do not accept checks or money orders. Purchases may be returned within 14 days of receipt for refund; refund will be initiated upon receipt of returned book.
Goodwill Books
2920 SE Century Blvd
Hillsboro, OR 97123
gicwbook@gicw.org
Shipping terms
Shipping costs are based on books weighing 2.2 LB, or 1 KG. If your book order is heavy or oversized, we may contact you to let you know extra shipping is required.