Items related to Sculpting Data for ML: The first act of Machine Learning

Sculpting Data for ML: The first act of Machine Learning - Softcover

Grover, Jigyasa; Misra, Rishabh

  • 4.80 out of 5 stars
    5 ratings by Goodreads
 
9798585463570: Sculpting Data for ML: The first act of Machine Learning

Synopsis

In the contemporary world of Artificial Intelligence and Machine Learning, data is the new oil. For Machine Learning algorithms to work their magic, it is imperative to lay a firm foundation with relevant data. Sculpting Data for ML introduces the readers to the first act of Machine Learning, Dataset Curation. This book puts forward practical tips to identify valuable information from the extensive amount of crude data available at our fingertips. The step-by-step guide accompanies code examples in Python from the extraction of real-world datasets and illustrates ways to hone the skills of extracting meaningful datasets. In addition, the book also dives deep into how data fits into the Machine Learning ecosystem and tries to highlight the impact good quality data can have on the Machine Learning system's performance.

What's Inside?
* Significance of data in Machine Learning
* Identification of relevant data signals
* End-to-end process of data collection and dataset construction
* Overview of extraction tools like BeautifulSoup and Selenium
* Step-by-step guide with Python code examples of real-world use cases
* Synopsis of Data Preprocessing and Feature Engineering techniques
* Introduction to Machine Learning paradigms from a data perspective

This book is for Machine Learning researchers, practitioners, or enthusiasts who want to tackle the data availability challenges to address real-world problems.

The authors Jigyasa Grover & Rishabh Misra are Machine Learning Engineers by profession and are passionate about tackling real-world problems leveraging their data curation and ML expertise.

The book is endorsed by leading ML experts from both academia and industry. It has forewords by:
* Julian McAuley, Associate Professor at University of California San Diego
* Laurence Moroney, Lead Artificial Intelligence Advocate at Google
* Mengting Wan, Senior Applied Scientist at Microsoft

"synopsis" may belong to another edition of this title.

About the Author

                   
Jigyasa Grover
Jigyasa Grover is a Machine Learning Engineer at Twitter. She has a myriad of experiences from her brief stints at Facebook, National Research Council of Canada, and Institute of Research & Development France involving Data Science, mathematical modeling, and software engineering. Having graduated from the UC San Diego, with a Master's degree in Computer Science, she is presently plying her past experiences & knowledge towards Applied ML in the online advertisements prediction & ranking domain. Red Hat 'Women in Open Source' Academic Award Winner and Google Summer of Code alumna, Jigyasa is an ardent open-source contributor as well. She served as the Director of Women Who Code & Lead of Women Techmakers for a handful of years to help bridge the gender gap in technology. In her quest to build a powerful community of girls and boys alike, and believing in "we rise by lifting others," she mentors aspiring developers and Machine Learning enthusiasts in various global programs.

Rishabh Misra
Rishabh Misra is a Machine Learning Engineer at Twitter. He developed a passion for identifying & tackling novel and practical problems using ML during his research internships at the Indian Institute of Technology Madras, which he further explored during his Master's in Computer Science from the UC San Diego. He combines his past engineering experiences in designing large-scale systems, working at Amazon & Arcesium, and research experiences in Applied Machine Learning to develop distributed Machine Learning relevance systems at Twitter. His explorations have led to several research publications in competitive ML conferences like RecSys, ACL, and WSDM. The ML community has well received the datasets collected (also used in this book) as part of his research. Kaggle recently ranked him as one of the top 20 dataset contributors, & Deeplearning.ai's "Natural Language Processing in TensorFlow" course on Coursera used his Sarcasm Detection dataset for illustrations.
            

From the Back Cover

In the contemporary world of Artificial Intelligence and Machine Learning, data is the new oil. For Machine Learning algorithms to work their magic, it is imperative to lay a firm foundation with relevant data. Sculpting Data for ML introduces the readers to the first act of Machine Learning, Dataset Curation. This book puts forward practical tips to identify valuable information from the extensive amount of crude data available at our fingertips. The step-by-step guide accompanies code examples in Python from the extraction of real-world datasets and illustrates ways to hone the skills of extracting meaningful datasets. In addition, the book also dives deep into how data fits into the Machine Learning ecosystem and tries to highlight the impact good quality data can have on the Machine Learning system's performance.

What's Inside?

  • Significance of data in Machine Learning
  • Identification of relevant data signals
  • End-to-end process of data collection and dataset construction
  • Overview of extraction tools like BeautifulSoup and Selenium
  • Step-by-step guide with Python code examples of real-world use cases
  • Synopsis of Data Preprocessing and Feature Engineering techniques
  • Introduction to Machine Learning paradigms from a data perspective

This book is for Machine Learning researchers, practitioners, or enthusiasts who want to tackle the data availability challenges to address real-world problems.

Jigyasa Grover & Rishabh Misra are Machine Learning Engineers by profession and are passionate about tackling real-world problems leveraging their data curation and ML expertise.

"About this title" may belong to another edition of this title.