Research Ideas and Outcomes : Software Description
PDF
Software Description
A Browser-Based Curation Tool for Expert Review of DNA Barcode Records from BOLD Systems
expand article infoStephan Kühbandner, Fabian Deister, Torbjørn Ekrem§, Ben Price|, Elisabeth Stur§, Brent Emerson, Peter M. Hollingsworth#, Rutger Aldo Vos¤, Michael J. Raupach, Leonardo Dapporto«, Adele Bordoni«, Claudia Bruschini«, Sónia Ferreira»,˄, Axel Hausmann
‡ Bavarian State Collection of Zoology, Munich, Germany
§ Department of Natural History, NTNU University Museum, Norwegian University of Science and Technology, Trondheim, Norway
| Natural History Museum, London, United Kingdom
¶ Instituto de Productos Naturales y Agrobiología (IPNA-CSIC), La Laguna, Tenerife, Islas Canarias, Spain
# Royal Botanic Garden Edinburgh, Edinburgh, United Kingdom
¤ Naturalis Biodiversity Center, Leiden, Netherlands
« Dipartimento di Biologia, Università degli Studi di Firenze, Florence, Italy
» CIBIO, Centro de Investigação em Biodiversidade e Recursos Genéticos, InBIO Laboratório Associado, Campus de Vairão, Universidade do Porto, 4485-661 Vairão, Vila do Conde, Portugal
˄ BIOPOLIS Program in Genomics, Biodiversity and Land Planning, CIBIO, Campus de Vairão, 4485-661 Vairão, Vila do Conde, Portugal
Open Access

Abstract

Background

We present a browser-based curation tool (Library Curation Tool) developed to support expert validation of taxonomic records derived from the Barcode of Life Data System (BOLD). This tool forms a critical component of a two-step approach designed within the EU Horizon Europe project Biodiversity Genomics Europe (BGE) to build a high-quality, curated DNA barcode reference library for European species. The upstream component—a bioinformatics pipeline described in a companion publication—automatically filters, cleans, and ranks BOLD records based on metadata completeness, sequence quality, and taxonomic consistency. However, certain complex cases, such as misidentifications, nomenclatorial problems (e.g. synonymy), BIN-sharing (multiple species sharing one BIN) or BIN-splitting (a single species associated with multiple BINs), cannot be fully resolved by automated methods and require expert judgment.

New information

Our Library Curation Tool enables taxonomic experts to interactively inspect, validate, or exclude individual records, update species names, assign curation statuses, and provide curator notes. The tool supports real-time statistics for BIN conflicts and dynamically updates curation metrics as the expert interacts with the data. Its user interface is designed to simplify the review of large datasets while ensuring consistency, traceability, and minimal risk of structural errors common in spreadsheet-based curation workflows.

The curated output from this tool, combined with the automated pipeline, forms the foundation of a reference library suitable for accurate DNA-based species identification in biodiversity monitoring and ecological studies. By integrating expert knowledge into a standardized and scalable interface, the tool supports distributed community curation of DNA barcode reference data. Although currently implemented as a local application, the workflow is designed to facilitate the consolidation of expert annotations into shared, FAIR-compliant reference libraries and future integration with community infrastructures such as BOLD and BOLD-Europe.

Keywords

Reference library curation, BOLD, Taxonomic records, BIN, DNA barcoding

Introduction

Background

DNA barcoding (Hebert et al. 2003, DeSalle and Goldstein 2019, Rani et al. 2026) has become a widely used approach for species identification and biodiversity assessment across a broad range of taxa. The Barcode of Life Data System (BOLD, Ratnasingham and Hebert 2007) is the central repository for animal DNA barcode data, hosting over 17.8 million specimen records globally from 1.3 million Barcode Index Numbers (BINs), which are proxies for species (Ratnasingham and Hebert 2013,Ratnasingham 2024a), including over 1.5 million records for European species alone. Despite the enormous value of this resource, many of the records — particularly those from early barcoding efforts or mirrored from GenBank — often lack critical metadata, contain outdated taxonomic names, or do not meet current quality standards (Baena-Bejarano et al. 2023,Radulovici et al. 2021, Lavrador et al. 2023). As a result, accurate species-level identification using BOLD data often depends on extensive post-processing and expert validation.

To address these limitations, the EU Horizon Europe project Biodiversity Genomics Europe (BGE, Naturalis Biodiversity Center 2025) is developing a curated DNA barcode reference library for European species, with a primary focus on pollinators, freshwater, and marine taxa. The reference library curation is performed in two steps (Fig. 1). In the first phase, an automated bioinformatics pipeline (Vos 2024, Price 2025) filters and ranks public BOLD records (Ratnasingham 2024b) based on metadata completeness, sequence quality, taxonomic consistency, and known issues such as synonymy and typographical errors. This process results in a pre-curated dataset suitable for further review.

Figure 1.  

The library curation workflow starts by mining data from BOLD, then processes this data with an automated curation pipeline, before following up with manual curation by taxonomic experts using the Library Curation Tool and publishing the curated reference library as a dataset on BOLD.

However, automated filtering alone is insufficient to resolve certain biologically complex or taxonomically ambiguous cases. For example, BIN-sharing events (multiple species share a single Barcode Index Number) or BIN-splitting events (a single species is assigned to multiple BINs) require expert taxonomic knowledge to interpret and resolve (Fontes et al. 2021, Hausmann et al. 2013). Furthermore, taxonomy is a dynamic discipline in which species concepts, nomenclature, and systematic relationships are continuously revised as new evidence becomes available. Consequently, DNA barcode reference libraries cannot be regarded as static resources but require ongoing review and maintenance by distributed taxonomic experts. Ensuring both the quality and long-term relevance of reference databases therefore depends on coordinated community curation efforts. The workflow presented here addresses this challenge by combining automated pre-curation with expert-driven review in a standardized environment, enabling taxonomic expertise to be captured, documented, and incorporated into reference library development.

Several approaches have been proposed to support the systematic curation of BOLD records. One such approach is the Barcode Audit and Grade System (BAGS; Fontes et al. 2021), which assigns grades (A–E) to species based on the number of records per BIN and the occurrence of BIN-sharing or BIN-splitting (Table 1). Within the BGE project, this concept has been extended through the definition of country representatives — pre-selected records with the best metadata quality for each combination of species, OTU and country. OTUs are used in addition to species and BIN assignments to ensure that geographically distributed genetic variation within species is represented during representative selection, while the country-based approach prevents over-reliance on records from a limited number of heavily sampled regions. Together, these criteria promote both geographic and genetic representation within the curated reference library.

Table 1.

BAGS - Barcode, Audit & Grade System.

GRADE
A >10 specimens in 1 BIN
B 3-10 specimens in 1 BIN
C >1 BIN
D <3 specimens in 1 BIN
E BIN sharing (>1 species in single BIN)

Additionally, we have developed a metadata quality rating system (Price 2025, Vos et al. 2026) that ranks records on a scale from 1 to 6, incorporating factors such as sequence length, presence of voucher information, and completeness of collection data (Table 2). These frameworks provide the quantitative foundation for prioritizing records during manual review.

Table 2.

Ranking system to pick representatives for each haplotype / species: Ranking 1-3 means records with good metadata quality (highlighted in grey), which will be pre-selected for the reference library and ranking 4-6 records with bad metadata quality, which are not pre-selected for the reference library. For "Public voucher" "✔(or)" means that only one of theses criteria has to be fulfilled in order to meet the ranking for all criteria with this prefix. The same is true for "Collection".

specimen rank
Criteria 1 2 3 4 5 6
Species level ID
Type specimen
Good quality sequence
Image(s) available
Collection country
ID identifier named
ID method (method is not BIN match)
Public voucher (has museum ID) ✔(or) ✔(or)
Public voucher (agreed institution) ✔(or) ✔(or)
Public voucher (agreed voucher type) ✔(or) ✔(or)
Collection (Date)
Collection (Site) ✔(or)
Collection (GPS coordinate) ✔(or)
Collection (Sector) ✔(or)
Collection (Region) ✔(or)
Collector named

While it is technically possible to conduct manual curation using common spreadsheets, this approach becomes impractical and error-prone for large and metadata-rich datasets (Broman and Woo 2018). Spreadsheet-based workflows are prone to formatting inconsistencies, accidental overwriting of fields, unstandardized status entries, and loss of data integrity when files are transferred between curators (Broman and Woo 2018). They also lack the ability to dynamically update key indicators such as BIN-sharing, BIN-splitting, or BAGS scores in real time. In contrast, a dedicated curation tool can enforce consistent data structures, provide immediate feedback on changes, and reduce the volume of data that needs to be exchanged with experts (Broman and Woo 2018).

To support this critical second phase, we developed a dedicated, browser-based curation tool tailored to the needs of taxonomic experts.

Here we introduce the design and functionality of the Library Curation Tool, and highlight its potential role in producing high-quality barcode reference data for DNA-based species identification. In doing so, we aim to provide a scalable, transparent, and expert-driven solution for curating large and complex barcode datasets, particularly in the context of biodiversity research and monitoring initiatives.

Project description

Title: 

BGE Library Curation Tool

Design description: 

The Libary Curation Tool allows curators to review, validate, and annotate BOLD-derived records using a structured, user-friendly interface (cf. Fig. 2). It provides real-time statistics on species coverage, BIN conflicts, and curation progress, and ensures data integrity through constrained input options and exportable, version-ready outputs.

Figure 2.  

Main Interface of the Library Curation Tool with labels explaining the different sections.

The manual curation process using this tool generally follows these steps:

  1. Load dataset: Select the .db file containing the data for the target taxonomic group generated by the pre-curation pipeline.

  2. Filter and search: Narrow down the dataset by species, BIN, or other metadata fields.

  3. Review records: Examine species names, BIN assignments, metadata quality, and potential conflicts.

  4. Assign status: Mark each record as valid, invalid, or excluded; correct species names where needed.

  5. Add notes: Document curation decisions with curator comments.

  6. Monitor statistics: Use dynamic counters to track BIN-sharing/splitting events and curation completeness.

  7. Export results: Save curated data as .csv along with a changes.log file for audit purposes.

Funding: 

Biodiversity Genomics Europe (Grant no.101059492) is funded by Horizon Europe (European Comission 2021) under the Biodiversity, Circular Economy and Environment call (REA.B.3); co-funded by the Swiss State Secretariat for Education, Research and Innovation (SERI) under contract numbers 22.00173 and 24.00054; and by the UK Research and Innovation (UKRI) under the Department for Business, Energy and Industrial Strategy’s Horizon Europe Guarantee Scheme.

Web location (URIs)

Technical specification

Platform: 
Browser (Edge, Chrome, Firefox, etc.)
Programming language: 
Java Script, HTML, CSS
Operational system: 
Windows, Linux, Mac OS
Interface language: 
English

Repository

Type: 
Git

Usage licence

Usage licence: 
Creative Commons Public Domain Waiver (CC-Zero)

Implementation

Implements specification

The Library Curation Tool (cf. Fig. 3) is a browser-based application designed to assist taxonomic experts in the manual validation of DNA barcode records, particularly those derived from the Barcode of Life Data System (BOLD). It serves as the expert-driven interface in a two-phase curation workflow. The upstream component is a semi-automated bioinformatics pipeline that filters, ranks, and enriches raw BOLD data. This tool builds on that output by providing an intuitive environment for manual review and expert decision-making. The steps of manual curation carried out by taxonomic experts with this tool comprise inspection of records (view pre-selected records, filter data, view BAGS grade, BIN-sharing and -splitting events and metadata) and take action (validate or invalidate records, exclude (reinclude) species, change species name, choose reason for name corrections and add curator notes as freetext).

Figure 3.  

HTML table of the Library Curation Tool with taxonomic records from BOLD and additional curation specific metadata (like: url, Ranking, country_representative, BAGS, Status, Reason Name Correction, Correct Species Name, Curator Notes). Records with grey background are pre-selected for reference library and need no action by the user for getting them added to the reference library. However, users can validate non-pre-selected records, which will get a green background or invalidate pre-selected records, which will have a red background.

All actions performed in the Library Curation Tool are documented in log files including timestamps, the executed action, and all information provided by the experts. These log files are analysed using a dedicated script (https://github.com/FabianDeister/BGE_library_curation_tool_log_processing) together with the output of the automated pipeline, whereby timestamps ensure that only the most recent version of each change is retained. In this way, both automatically pre-validated and manually reviewed records are merged to form the curated reference library. The resulting curated datasets provide the basis for the development of curated European DNA barcode reference libraries within the Biodiversity Genomics Europe project. In the longer term, the project aims to make curated outputs and expert annotations accessible through shared infrastructures, including BOLD and BOLD-Europe, allowing curation decisions to contribute to community-maintained reference resources. The exact mechanisms for integration and publication are currently under discussion with BOLD partners.

The application is implemented using standard web technologies and can be run entirely on a local computer, without internet access (aside from optional loading of remote CSS assets). It includes the following components:

  • Backend: The backend logic is handled by a lightweight Node.js (Open JS Foundation 2024) server (server.js), which manages HTTP requests and interacts with the input dataset—a structured SQLite-compatible .db file produced by the upstream curation pipeline. This file contains metadata-rich sequence records, including taxonomic information, BIN URIs, quality scores, and precomputed BAGS values.

  • Frontend: The user interface is built in index.html using HTML, CSS, and JavaScript, and runs in a modern web browser (e.g. Chrome, Firefox). It utilizes the DataTables library to provide interactive, searchable, and paginated tables. Custom JavaScript code supports advanced functionalities such as row coloring based on status, BIN visualization, in-table dropdowns for status selection, curator note entry, and per-record submission.

  • Database Input: The tool operates on local .db database files. These files are placed in the data/ subdirectory and loaded dynamically through a dataset selector. Each file corresponds to a taxonomic group and contains hundreds to thousands of records to be curated.

  • Execution Environment: The application is platform-independent and distributed as a self-contained folder Fig. 4. On Windows, users simply double-click start_tool.bat to launch the server and automatically open the tool in a browser via http://localhost:3000. On Linux and macOS, the tool can be launched manually from the command line using Node.js. A detailed installation manual for Linux and macOS can be found here: https://github.com/bge-barcoding/BGE_library_curation_tool.

  • Export and Audit: All curation actions (status changes, species name updates, notes) are logged in a changes.log file, ensuring transparency and reproducibility. Curators can export their results (for own purposes) in .csv format and submit the log file as a standardized feedback mechanism. The tool prevents structural errors common in spreadsheet-based curation by enforcing consistent fields and controlled input types.

  • Dynamic Scoring: The tool includes dynamic logic for recalculating BAGS scores and BIN statistics in real time. This allows experts to see how their actions (e.g. excluding a species or marking a record invalid) influence BIN-sharing, BIN-splitting, and representative selection.

Figure 4.  

Curation Tool - main folder.

In summary, the Library Curation Tool is a locally hosted, browser-accessible interface purpose-built for scalable expert curation of DNA barcode data. It bridges the gap between automated pipeline output and final expert-reviewed reference libraries, facilitating the creation of high-quality, FAIR-compliant resources for molecular biodiversity research.

There are several levels of support for the user. First, there is a user manual within the main folder of the curation tool. Second, contextual help is provided the user interface by red question marks - clicking on them opens a menu with additional information. Third, a video tutorial and FAQ section are available on the project website (https://bge-barcoding.github.io/manual-curation/). An overview of the associated github repositories is presented in Table 3.

Although the current implementation operates locally on the curator's computer, the workflow is designed around standardized data structures, controlled vocabularies, reproducible log files, and version-controlled outputs. These features facilitate the transparent exchange of curation decisions and support future integration into shared reference data infrastructures. Rather than promoting isolated local reference databases, the long-term objective is to enable expert contributions from distributed specialists to be consolidated into community-curated, FAIR-compliant reference libraries that are accessible and reusable across projects, institutions, and countries.

Audience

Taxonomic experts will curate records from BOLD Systems that have been pre-curated using this pipeline: https://github.com/bge-barcoding/bold-library-curation.

References

login to comment