Full Paper

Automatic Metadata Extraction From Museum Specimen Labels

  • P. Bryan Heidorn
  • Qin Wei
  • University of Illinois at Urbana-Champaign
Published
Access and licence
Open access CC BY 4.0
Download PDF

Abstract

This paper describes the information properties of museum specimen labels and machine learning tools to automatically extract Darwin Core (DwC) and other metadata from these labels processed through Optical Character Recognition (OCR). The DwC is a metadata profile describing the core set of access points for search and retrieval of natural history collections and observation databases. Using the HERBIS Learning System (HLS) we extract 74 independent elements from these labels. The automated text extraction tools are provided as a web service so that users can reference digital images of specimens and receive back an extended Darwin Core XML representation of the content of the label. This automated extraction task is made more difficult by the high variability of museum label formats, OCR errors and the open class nature of some elements. In this paper we introduce our overall system architecture, and variability robust solutions including, the application of Hidden Markov and Naïve Bayes machine learning models, data cleaning, use of field element identifiers, and specialist learning models. The techniques developed here could be adapted to any metadata extraction situation with noisy text and weakly ordered elements.

The full text of this article is available as a PDF.

Download PDF

Article details

Published
Section
Full Papers
DOI
10.23106/dcmi.952109189
License
CC BY 4.0 · open access

Indexed in

Described in Dublin Core

This article's metadata, in the vocabulary these proceedings are about.

dcterms:title
Automatic Metadata Extraction From Museum Specimen Labels
dcterms:creator
Heidorn, P. Bryan
Wei, Qin
dcterms:date
2008-09-10
dcterms:identifier
doi:10.23106/dcmi.952109189
dcterms:subject
automatic metadata extraction
machine learning
Hidden Markov Model
Naïve Bayes
Darwin Core
dcterms:publisher
Dublin Core Metadata Initiative
dcterms:type
Text
dcterms:language
en
dcterms:rights
CC BY 4.0