Constructing biological knowledge bases by extracting information from text sources.

Mark Craven; Johan Kumlien

Constructing biological knowledge bases by extracting information from text sources.

Mark Craven(Carnegie Mellon University), Johan Kumlien

PubMed

January 1, 1999

Cited by 597

Abstract

Recently, there has been much effort in making databases for molecular biology more accessible and interoperable. However, information in text form, such as MEDLINE records, remains a greatly underutilized source of biological information. We have begun a research effort aimed at automatically mapping information from text sources into structured representations, such as knowledge bases. Our approach to this task is to use machine-learning methods to induce routines for extracting facts from text. We describe two learning methods that we have applied to this task--a statistical text classification method, and a relational learning method--and our initial experiments in learning such information-extraction routines. We also present an approach to decreasing the cost of learning information-extraction routines by learning from "weakly" labeled training data.

Judea Pearl|Unknown|1988|17k

On the Optimality of the Simple Bayesian Classifier under Zero-One Loss

Pedro Domingos, Michael J. Pazzani|Machine Learning|1997|3.1k

[31] PHD: Predicting one-dimensional protein structure by profile-based neural networks

Burkhard Rost|Methods in enzymology on CD-ROM/Methods in enzymology|1996|1.3k

Learning Information Extraction Rules for Semi-Structured and Free Text

Stephen Soderland|Machine Learning|1999|933

Information extraction

Jim Cowie, Wendy G. Lehnert|Communications of the ACM|1996|731

Constructing biological knowledge bases by extracting information from text sources.

Abstract

Related Papers