The BKH Services
Section outline
-
Description
This course presents the BKH services. Video Tutorials, Training Materials, and use cases are presented, offering a full overview of the biodiversity data workflow and the data transfer mechanisms available.
Target Audience
- Academic researchers
- Graduate students
- Journal editors and reviewers
- Scientific writers
Objectives
By the end of this course, participants will be able to:
- Explore the BKH Services and get familiarized with their use.
- Get an introduction to the specific functionalities of each service.
- Introduce the use of the BKH services to their everyday practices.
Contents
- Video Tutorials and Guides
- Access to the BKH Services
- Practical Examples of Use
-

Before we proceed to the description of the different services offered from BKH we present here in this section some basic information about the Transfer Formats and the Data Exchange Mechanisms used in BKH.
Following the presentation of these basic concepts we will present how the BKH facilitates the overall Workflow of the Biodiversity Data. In addition, we also introduce the tools that are forming the BKH Workbench.
The core part of this course will then focus on the presentation of BKH Services.
-
COL and GBIF have united their capabilities to make ChecklistBank, a publishing platform and repository for all taxonomic checklist datasets, nomenclatural datasets, and any other publicly published species lists. Through Plazi’s TreatmentBank tens of thousands of datasets from liberated publications are made available. ChecklistBank also contains classifications, species hypotheses, OTUs, and BINs from Barcode of Life, NCBI Taxonomy/ENA, and UNITE/PlutoF, amongst others. Data in ChecklistBank are used to assemble the Catalogue of Life Checklist to create a consistent and up-to-date listing of all the world’s known species. Data and tooling are also used to create custom taxonomic data products, such as the backbone taxonomy for GBIF. The COL Checklist has stable name usage identifiers, and the annual or monthly versions of the COL Checklist are issued with Digital Object Identifiers, also for the contributing data sources.
ChecklistBank is open for others to add and use. It generates a standardized interpretation of data. Tooling is offered to search for name usage across different datasets (checklists) and to compare name usage between datasets. All datasets can be searched, browsed, downloaded, or accessed programmatically via the ChecklistBank API.
To use all functions of ChecklistBanks you will need to log in with a GBIF user account. This includes a continuously improved diff tool for comparing datasets as well as downloading custom exports.
-

The e-Biodiversity matching services, in short e-Biodiv, are a set of services and Graphic User Interfaces to support the linking of specimen and material citation data. These pairs are based on material citations in GBIF mediated by Plazi and discovered by applying the clustering algorithm in GBIF. -
LifeBlock is a service developed by LifeWatch ERIC that provides federated search and semantic search on any of the research infrastructures. The user can search and then store the results in a dedicated account that preserves the traceability of all the information management following FAIR principles. LifeBlock can search both metadata and data. Additionally, advanced users can create their own applications and implement them to manage the information retrieved.
-

OpenBiodiv offers a broad biodiversity-related querying system answering open-ended queries based on the data extracted and converted to RDF from Pensoft journals and Plazi’s taxon treatments. Data can be explored in four different ways: General search, SPARQL, User applications and API. OpenBiodiv can discover hidden links within biodiversity data (e.g. between authors, taxa, sequences, material citations, publications, and others) and can guide research into how data is used in scholarly articles.
-
PlutoF offers a third-party curation service to improve the quality of public DNA sequences and their source metadata (e.g., on material source, geolocation and habitat, taxonomic identifications, interacting taxa, literature, etc.) In collaboration with EMBL-EBI, improved or corrected annotations on INSD sequences in PlutoF are fed back to primary repositories through operating the ELIXIR Contextual Data ClearingHouse. Searching and browsing of third-party annotations introduced by the UNITE Community can be done via RESTful API or using search interfaces of PlutoF and UNITE.
-

Biodiversity PMC is a custom search engine to explore biodiversity literature alongside biomedical literature (e.g., PMC, MEDLINE). The service provides access to Pazi’s taxonomic treatments in JATS/TaxPub formats. Based on the SIB Literature Services (SIBiLS), all contents have been semantically enriched with several domain-specific onto-terminologies (e.g. NCBI Taxonomy, Open Tree of Life, Gene Ontology) to enable users to author semantically enhanced queries. Inspired by the “One Health” recommendations of WHO, the service is enhancing the coverage of PubMed Central by integrating a growing collection of open-access full-text publications in biodiversity-related fields, such as ecology and taxonomy. The data are distributed under CC-BY licenses in BioC formats.
-
The biotic interactions browser searches the literature to identify pairs of species that could have a biotic relationship. Each identified interaction is accompanied by one or several relevant passages extracted from the scientific literature. With this tool, users can discover new biotic interactions and understand how they are established.
-

The SIBiLS SPARQL endpoint allows to running of SPARQL queries on a subset of annotated PMC publications. The textual content and structure (title, sections, paragraphs, etc.) are described as well as a list of named entities (multidimensional concepts) found and located in the full text of the publications. The ontology used to describe the publications can be found here. The triple store contains a sample of 3000 publications at the moment. The content of the triple store will be extended in the future and include alternate publication sources as well.
-
Synospecies uses published taxonomic treatment citations to represent the history of taxonomic names. A visualization provides an overview of the history and a SPARQL endpoint is available for exploration. The data is provided by TreatmentBank and is available as RDF.
-


TreatmentBank provides an access point to explore and discover data liberated from publications, and to bidirectional links created by the deposition of data to respective services (Catalogue of Life / ChecklistBank; GBIF; ENA; Zenodo (Biodiversity Literature Service). -

The Biodiversity Literature Repository (BLR) is a research infrastructure (RI) comprising the BLR Community on Zenodo at the European Center for Nuclear Research (CERN), and services to search and retrieve the data (Ocellus, Zenodeo API, BLR website). BLR’s focus is on biodiversity data liberated from scholarly publications, and it uses custom metadata linking to external vocabularies covering the needs of the biodiversity community. This includes taxonomic treatment or figures as well as the deposit of the original article annotated with metadata describing the data contained in the articles itself, as well as related identifiers for figures and treatments therein. The main data import is through TreatmentBank or via publishers such as Pensoft. With over 650,000 deposits, BLR is the single largest community in Zenodo. Its data is widely reused, for example by the Global Biodiversity Information Facility (GBIF). All data in BLR is published under the CC0 Public Domain Dedication, remaining free for anyone to use, anywhere, for any purpose.
BLR’s role in BiCIKL is to provide long-term virtual access to data liberated from scientific literature, as well as to data liberated through transnational access.
-

ARPHA Writing Tool 2.0 (AWT 2.0) is an XML-based, WYSIWYG (What-You-See-Is-What-You-Get) authoring tool, which allows co-authors and collaborators to work conveniently together on manuscripts before submission to a journal. To ensure interoperability and prompt re-use of data after publication, the tool relies on semantic enhancements and PIDs, vocabularies, and ontologies.
To cater to the specifics of biodiversity publications, AWT 2.0 supports various domain-specific workflows, including data import and export, as well as bi-directional links with leading aggregators, e.g. GBIF, BOLD, INSDC, Catalogue of Life, ChecklistBank, TreatmentBank, BHL, Biodiversity Literature Repository, BiodiversityPMC (SiBILS) and others.
-
Nanopublications allow researchers to ‘fragment’ their most important scientific findings into ‘pixels of knowledge’, where each assertion becomes findable, accessible, interoperable, and reusable (FAIR). Nanopublications can also be used to annotate existing publications or other online resources.
The Nanopublications-for-Biodiversity workflow and templates support various associations, such as those between organisms, between taxa, between taxa and environments, and between organisms and nucleotide sequences amongst others. To do this, the domain-specific workflow and templates rely on community-agreed and widely used standards and persistent identifiers and API services from the likes of ChecklistBank, Catalogue of Life, GBIF, GenBank/ENA, BOLD, Darwin Core, ZooBank, Index Fungorum, MycoBank, IPNI, and TreatmentBank.
Nanopublications can be published in association with a manuscript in BDJ, as standalone publications related to any other article or online resource, or as annotations to any article published on ARPHA.
Nanopublications-for-Biodiversity is a collaboration project between Pensoft and Knowledge Pixels AG using the Nanodash tool.