Automatic concept recognition using the Human Phenotype Ontology reference and test suite corpora

T. Groza; S. Kohler; S. Doelken; N. Collier; A. Oellrich; D. Smedley; F.M. Couto; G. Baynam; A. Zankl; P.N. Robinson

doi:10.1093/database/bav005

Back

Automatic concept recognition using the Human Phenotype Ontology reference and test suite corpora

Journal article

Open access

Peer reviewed

Automatic concept recognition using the Human Phenotype Ontology reference and test suite corpora

T. Groza, S. Kohler, S. Doelken, N. Collier, A. Oellrich, D. Smedley, F.M. Couto, G. Baynam, A. Zankl and P.N. Robinson

Database, Vol.2015, pp.1-13

2015

DOI: https://doi.org/10.1093/database/bav005

Files and links (2)

pdf

Database-2015-Groza-database_bav005.pdfDownload View

Published (Version of Record) Open Access

url

Free to Read *No subscription requiredView

Abstract

Concept recognition tools rely on the availability of textual corpora to assess their performance and enable the identification of areas for improvement. Typically, corpora are developed for specific purposes, such as gene name recognition. Gene and protein name identification are longstanding goals of biomedical text mining, and therefore a number of different corpora exist. However, phenotypes only recently became an entity of interest for specialized concept recognition systems, and hardly any annotated text is available for performance testing and training. Here, we present a unique corpus, capturing text spans from 228 abstracts manually annotated with Human Phenotype Ontology (HPO) concepts and harmonized by three curators, which can be used as a reference standard for free text annotation of human phenotypes. Furthermore, we developed a test suite for standardized concept recognition error analysis, incorporating 32 different types of test cases corresponding to 2164 HPO concepts. Finally, three established phenotype concept recognizers (NCBO Annotator, OBO Annotator and Bio-LarK CR) were comprehensively evaluated, and results are reported against both the text corpus and the test suites. The gold standard and test suites corpora are available from http://bio-lark.org/hpo_res.html.

Details

Title: Automatic concept recognition using the Human Phenotype Ontology reference and test suite corpora
Authors/Creators: T. Groza (Author/Creator)
S. Kohler (Author/Creator)
S. Doelken (Author/Creator) - King's College London
N. Collier (Author/Creator)
A. Oellrich (Author/Creator) - Social Genetic & Developmental Psychiatry
D. Smedley (Author/Creator)
F.M. Couto (Author/Creator)
G. Baynam (Author/Creator)
A. Zankl (Author/Creator)
P.N. Robinson (Author/Creator)
Publication Details: Database, Vol.2015, pp.1-13
Publisher: Oxford University Press
Identifiers: 991005541536607891
Murdoch Affiliation: Institute for Immunology and Infectious Diseases
Language: English
Resource Type: Journal article

Metrics

270 File views/ downloads

59 Record Views