Semantic Tagging, Enrichment, and Data Extraction with PLAZI
Building on existing mechanisms and supported by the BiCIKL project, PLAZI has developed tools using semantic tagging and enrichment to extract data from the Global Biodiversity Information Facility (GBIF) and link them to open publications hosted on Zenodo. The integration ensures all outputs are citable and trackable via Digital Object Identifiers (DOIs).
This module guides you through three core stages of the data extraction workflow:
-
Preparing Data Extraction
- Understanding the linking process
- Getting familiar with GoldenGATE Imagine (GGI)
-
Performing Data Extraction
- Guide to individual paper extraction
- Querying repositories, analyzing stats, and assessing data reuse
- Utilizing the eBioDiv Matching Service
-
Improving Tool Results
- Quality Control (QC) checks
- Enhancing annotations
1. Preparing Data Extraction
Before extracting data from scientific literature, users must understand the underlying linking mechanisms and master the extraction software.
Understanding the Linking Process
- Semantic Tagging: Tags taxonomic names, bibliographic references, and treatment data with unique identifiers to ensure machine-readability.
- Persistent Identifiers: Assigns DOIs to taxonomic treatments published on PLAZI for permanent referencing.
- GBIF Taxonomic Backbone Integration: Connects names to a global, standardized taxonomy for cross-dataset interoperability.
- Cross-Referencing: Maps taxa to occurrence records and geographical distribution maps on GBIF.
- Information Enrichment: Enriches treatments with additional ecological and distributional context from global repositories.
- Direct Article Links: Provides embedded hyperlinks in PLAZI literature to external biodiversity databases.
- Data Accessibility & Exchange: Fosters open data exchange across global research infrastructures.
Getting Familiar with the Software
The data liberation workflow relies on GoldenGATE Imagine (GGI). To set up and learn basic operations, consult the official guide:
Tutorial: Getting used to GGI and its tools and functions (PDF)
2. Performing Data Extraction
Data extraction can be executed on individual publications, queried across platforms for analytical reporting, or processed through matching algorithms.
Individual Extraction Workflow
Review the GGI Glossary before beginning. Follow the step-by-step extraction workflow outlined in the Individual Extraction Guide (PDF):
- Detect Document Structure
- Add Document Metadata
- Parse Bibliography
- Mark Bibliographic Citations
- Mark Taxon Names
- Mark Taxon Keys
- Mark Treatments
- Define Treatment Structure
- Mark Treatment Citations
- Mark Materials Citations
- Parse Materials Citations
- Select Zenodo License
Queries, Data Repositories, and Reuse
Extracted treatments are distributed and queryable across several core platforms:
For analytics and statistical tracking (including Article, Treatment, and Treatment Details Linking stats), explore the Plazi Statistics Portal and consult the Repositories & Statistics Guide (PDF).
The eBioDiv Matching Service
The eBioDiv Matching Service connects material citations in literature with physical museum specimens through a semi-automated system. As human decisions refine candidate lists, high-confidence pairs will transition toward fully automated matching.
Resources for New Users
Resources for Advanced Users
3. Improving the Tool’s Results
Quality Control (QC)
The Quality Control system flags processing errors before data reaches public repositories. Error classifications range from critical blockers to minor warnings, covering Metadata, Text Streams, Treatments, Figures, Material Citations, Bibliographic References, and Taxon Names.
Enhancing Annotations
To maximize annotation precision, users can re-open and refine parsed data using built-in enhancement tools:
- Edit or Copy Annotation Attributes
- Parse Materials Citations, References, and Taxon Names
- Assign Captions to Figures and Tables
- Search or List Annotations and Words
- Revise Block Paragraph Structures