Cleaning Data for Effective Data Science: Doing the other 80% of the work with Python, R, and command-line tools

Data cleaning is the all-important first step to successful data science, data analysis, and machine learning. If you work with any kind of data, this book is your go-to resource, arming you with the insights and heuristics experienced data scientists had to learn the hard way.

In a light-hearted and engaging exploration of different tools, techniques, and datasets real and fictitious, Python veteran David Mertz teaches you the ins and outs of data preparation and the essential questions you should be asking of every piece of data you work with.

Using a mixture of Python, R, and common command-line tools, Cleaning Data for Effective Data Science follows the data cleaning pipeline from start to end, focusing on helping you understand the principles underlying each step of the process. You'll look at data ingestion of a vast range of tabular, hierarchical, and other data formats, impute missing values, detect unreliable data and statistical anomalies, and generate synthetic features. The long-form exercises at the end of each chapter let you get hands-on with the skills you've acquired along the way, also providing a valuable resource for academic courses.

Cleaning Data for Effective Data Science: Doing the other 80% of the work with Python, R, and command-line tools

29.99 In Stock

Cleaning Data for Effective Data Science: Doing the other 80% of the work with Python, R, and command-line tools

Add to Wishlist

Cleaning Data for Effective Data Science: Doing the other 80% of the work with Python, R, and command-line tools

eBook

$29.99

eBook
$29.99

Available on Compatible NOOK devices, the free NOOK App and in My Digital Library.

WANT A NOOK? Explore Now

Buy As Gift

Related collections and offers

Overview

Product Details

ISBN-13:	9781801074407
Publisher:	Packt Publishing
Publication date:	03/31/2021
Sold by:	Barnes & Noble
Format:	eBook
Pages:	498
File size:	9 MB

About the Author

David Mertz, Ph.D. is the founder of KDM Training, a partnership dedicated to educating developers and data scientists in machine learning and scientific computing. He created a data science training program for Anaconda Inc. and was a senior trainer for them. With the advent of deep neural networks, he has turned to training our robot overlords as well.

He previously worked for 8 years with D. E. Shaw Research and was also a Director of the Python Software Foundation for 6 years. David remains co-chair of its Trademarks Committee and Scientific Python Working Group. His columns, Charming Python and XML Matters, were once the most widely read articles in the Python world.

Table of Contents

Data Ingestion – Tabular Formats
Data Ingestion - Hierarchical Formats
Data Ingestion - Repurposing Data Sources
The Vicissitudes of Error - Anomaly Detection
The Vicissitudes of Error - Data Quality
Rectification and Creation - Value Imputation
Rectification and Creation - Feature Engineering
Ancillary Matters - Closure/Glossary

From the B&N Reads Blog

Page 1 of

Related collections and offers

Overview

Product Details

About the Author

Table of Contents

Related Subjects

Customer Reviews