Natural Language Annotation for Machine Learning
Language: English
Published by O'Reilly Media, US, 2012
- Softcover
- New

Seller: Rarewaves USA, HEBRON, KY, U.S.A.Rarewaves USA
AbeBooks seller since June 10, 2025
Condition: New
US$ 44.13
Quantity: Over 20 available
Add to basketItem description from seller
Create your own natural language training corpus for machine learning. This example-driven book walks you through the annotation cycle, from selecting an annotation task and creating the annotation specification to designing the guidelines, creating a "gold standard" corpus, and then beginning the actual data creation with the annotation process. Systems exist for analyzing existing corpora, but making a new corpus can be extremely complex. To help you build a foundation for your own machine learning goals, this easy-to-use guide includes case studies that demonstrate four different annotation tasks in detail. You'll also learn how to use a lightweight software package for annotating texts and adjudicating the annotations. This book is a perfect companion to O'Reilly's Natural Language Processing with Python, which describes how to use existing corpora with the Natural Language Toolkit.
Seller Inventory # LU-9781449306663
- Title
- Natural Language Annotation for Machine Learning
- Author
- James Pustejovsky
- Publisher
- O'Reilly Media, US
- Publication year
- 2012
- Condition
- New
- Binding
- Paperback
- Language
- English
- ISBN 10
- 1449306667
- ISBN 13
- 9781449306663
Create your own natural language training corpus for machine learning. Whether you’re working with English, Chinese, or any other natural language, this hands-on book guides you through a proven annotation development cycle―the process of adding metadata to your training corpus to help ML algorithms work more efficiently. You don’t need any programming or linguistics experience to get started.
Using detailed examples at every step, you’ll learn how the MATTER Annotation Development Process helps you Model, Annotate, Train, Test, Evaluate, and Revise your training corpus. You also get a complete walkthrough of a real-world annotation project.
- Define a clear annotation goal before collecting your dataset (corpus)
- Learn tools for analyzing the linguistic content of your corpus
- Build a model and specification for your annotation project
- Examine the different annotation formats, from basic XML to the Linguistic Annotation Framework
- Create a gold standard corpus that can be used to train and test ML algorithms
- Select the ML algorithms that will process your annotated data
- Evaluate the test results and revise your annotation task
- Learn how to use lightweight software for annotating texts and adjudicating the annotations
This book is a perfect companion to O’Reilly’s Natural Language Processing with Python.
"Synopsis" may belong to another edition of this title.
About the Author
Amber Stubbs recently completed her Ph.D. in Computer Science at Brandeis University, and is currently a Postdoctoral Associate at SUNY Albany. Her dissertation focused on creating an annotation methodology to aid in extracting high-level information from natural language files, particularly biomedical texts. Her website can be found at http://pages.cs.brandeis.edu/~astubbs/
"About the title" may belong to another edition of this title.
Shipping rates within U.S.A.
| Item | 10 to 13 business days | 10 to 13 business days |
|---|---|---|
| First item | US$ 0.00 | US$ 0.00 |
Payment methods
Seller's business information
Rarewaves USA
10100 West Sample Road, Ste 101
Coral Springs, FL U.S.A. 33065
Shipping terms
Please note that we do not offer Priority shipping to any country.
We currently do not ship to the below countries:
Afghanistan
Bhutan
Brazil
Brunei Darussalam
Channel Islands
Chile
Israel
Lao
Mexico
Russian Federation
Saudi Arabia
South Africa
Yemen
Please do not attempt to place orders with any of these countries as a ship to address - they will be cancelled.