Semantic Tagging, Enrichment, and Data Extraction with PLAZI

Building on existing mechanisms and supported by the BiCIKL project, PLAZI has developed tools using semantic tagging and enrichment to extract data from the Global Biodiversity Information Facility (GBIF) and link them to open publications hosted on Zenodo. The integration ensures all outputs are citable and trackable via Digital Object Identifiers (DOIs).

This module guides you through three core stages of the data extraction workflow:


1. Preparing Data Extraction

Before extracting data from scientific literature, users must understand the underlying linking mechanisms and master the extraction software.

Understanding the Linking Process

  • Semantic Tagging: Tags taxonomic names, bibliographic references, and treatment data with unique identifiers to ensure machine-readability.
  • Persistent Identifiers: Assigns DOIs to taxonomic treatments published on PLAZI for permanent referencing.
  • GBIF Taxonomic Backbone Integration: Connects names to a global, standardized taxonomy for cross-dataset interoperability.
  • Cross-Referencing: Maps taxa to occurrence records and geographical distribution maps on GBIF.
  • Information Enrichment: Enriches treatments with additional ecological and distributional context from global repositories.
  • Direct Article Links: Provides embedded hyperlinks in PLAZI literature to external biodiversity databases.
  • Data Accessibility & Exchange: Fosters open data exchange across global research infrastructures.

Getting Familiar with the Software

The data liberation workflow relies on GoldenGATE Imagine (GGI). To set up and learn basic operations, consult the official guide:

Tutorial: Getting used to GGI and its tools and functions (PDF)


2. Performing Data Extraction

Data extraction can be executed on individual publications, queried across platforms for analytical reporting, or processed through matching algorithms.

Individual Extraction Workflow

Review the GGI Glossary before beginning. Follow the step-by-step extraction workflow outlined in the Individual Extraction Guide (PDF):

  1. Detect Document Structure
  2. Add Document Metadata
  3. Parse Bibliography
  4. Mark Bibliographic Citations
  5. Mark Taxon Names
  6. Mark Taxon Keys
  7. Mark Treatments
  8. Define Treatment Structure
  9. Mark Treatment Citations
  10. Mark Materials Citations
  11. Parse Materials Citations
  12. Select Zenodo License

Queries, Data Repositories, and Reuse

Extracted treatments are distributed and queryable across several core platforms:

For analytics and statistical tracking (including Article, Treatment, and Treatment Details Linking stats), explore the Plazi Statistics Portal and consult the Repositories & Statistics Guide (PDF).

The eBioDiv Matching Service

The eBioDiv Matching Service connects material citations in literature with physical museum specimens through a semi-automated system. As human decisions refine candidate lists, high-confidence pairs will transition toward fully automated matching.

Resources for New Users
  1. Introduction to the Matching Service
  2. Data Source Overview
  3. Matching Algorithm Mechanics
Resources for Advanced Users
  1. Matching Service Overview (PDF)
  2. Getting Started Guide (PDF)
  3. Accessing the Matching Service (PDF)
  4. Matching Specimens with Material Citations (PDF)
  5. Matching Decisions (PDF)
  6. Matching Service Glossary

3. Improving the Tool’s Results

Quality Control (QC)

The Quality Control system flags processing errors before data reaches public repositories. Error classifications range from critical blockers to minor warnings, covering Metadata, Text Streams, Treatments, Figures, Material Citations, Bibliographic References, and Taxon Names.

Enhancing Annotations

To maximize annotation precision, users can re-open and refine parsed data using built-in enhancement tools:

  • Edit or Copy Annotation Attributes
  • Parse Materials Citations, References, and Taxon Names
  • Assign Captions to Figures and Tables
  • Search or List Annotations and Words
  • Revise Block Paragraph Structures