|
Research Ideas and Outcomes :
Software Description
|
|
Corresponding author: Stephan Kühbandner (kuehbandner@snsb.de), Torbjørn Ekrem (torbjorn.ekrem@ntnu.no)
Academic editor: Filipe Costa
Received: 17 Mar 2026 | Accepted: 20 Jun 2026 | Published: 14 Jul 2026
© 2026 Stephan Kühbandner, Fabian Deister, Torbjørn Ekrem, Ben Price, Elisabeth Stur, Brent Emerson, Peter Hollingsworth, Rutger Vos, Michael Raupach, Leonardo Dapporto, Adele Bordoni, Claudia Bruschini, Sónia Ferreira, Axel Hausmann
This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Citation:
Kühbandner S, Deister F, Ekrem T, Price B, Stur E, Emerson B, Hollingsworth PM, Vos R, Raupach MJ, Dapporto L, Bordoni A, Bruschini C, Ferreira S, Hausmann A (2026) A Browser-Based Curation Tool for Expert Review of DNA Barcode Records from BOLD Systems. Research Ideas and Outcomes 12: e191986. https://doi.org/10.3897/rio.12.e191986
|
|
We present a browser-based curation tool (Library Curation Tool) developed to support expert validation of taxonomic records derived from the Barcode of Life Data System (BOLD). This tool forms a critical component of a two-step approach designed within the EU Horizon Europe project Biodiversity Genomics Europe (BGE) to build a high-quality, curated DNA barcode reference library for European species. The upstream component—a bioinformatics pipeline described in a companion publication—automatically filters, cleans, and ranks BOLD records based on metadata completeness, sequence quality, and taxonomic consistency. However, certain complex cases, such as misidentifications, nomenclatorial problems (e.g. synonymy), BIN-sharing (multiple species sharing one BIN) or BIN-splitting (a single species associated with multiple BINs), cannot be fully resolved by automated methods and require expert judgment.
Our Library Curation Tool enables taxonomic experts to interactively inspect, validate, or exclude individual records, update species names, assign curation statuses, and provide curator notes. The tool supports real-time statistics for BIN conflicts and dynamically updates curation metrics as the expert interacts with the data. Its user interface is designed to simplify the review of large datasets while ensuring consistency, traceability, and minimal risk of structural errors common in spreadsheet-based curation workflows.
The curated output from this tool, combined with the automated pipeline, forms the foundation of a reference library suitable for accurate DNA-based species identification in biodiversity monitoring and ecological studies. By integrating expert knowledge into a standardized and scalable interface, the tool supports distributed community curation of DNA barcode reference data. Although currently implemented as a local application, the workflow is designed to facilitate the consolidation of expert annotations into shared, FAIR-compliant reference libraries and future integration with community infrastructures such as BOLD and BOLD-Europe.
Reference library curation, BOLD, Taxonomic records, BIN, DNA barcoding
DNA barcoding (
To address these limitations, the EU Horizon Europe project Biodiversity Genomics Europe (BGE,
However, automated filtering alone is insufficient to resolve certain biologically complex or taxonomically ambiguous cases. For example, BIN-sharing events (multiple species share a single Barcode Index Number) or BIN-splitting events (a single species is assigned to multiple BINs) require expert taxonomic knowledge to interpret and resolve (
Several approaches have been proposed to support the systematic curation of BOLD records. One such approach is the Barcode Audit and Grade System (BAGS;
| GRADE | |
| A | >10 specimens in 1 BIN |
| B | 3-10 specimens in 1 BIN |
| C | >1 BIN |
| D | <3 specimens in 1 BIN |
| E | BIN sharing (>1 species in single BIN) |
Additionally, we have developed a metadata quality rating system (
Ranking system to pick representatives for each haplotype / species: Ranking 1-3 means records with good metadata quality (highlighted in grey), which will be pre-selected for the reference library and ranking 4-6 records with bad metadata quality, which are not pre-selected for the reference library. For "Public voucher" "✔(or)" means that only one of theses criteria has to be fulfilled in order to meet the ranking for all criteria with this prefix. The same is true for "Collection".
| specimen rank | ||||||
| Criteria | 1 | 2 | 3 | 4 | 5 | 6 |
| Species level ID | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Type specimen | ✔ | |||||
| Good quality sequence | ✔ | ✔ | ✔ | ✔ | ||
| Image(s) available | ✔ | |||||
| Collection country | ✔ | ✔ | ✔ | |||
| ID identifier named | ✔ | ✔ | ||||
| ID method (method is not BIN match) | ✔ | ✔ | ||||
| Public voucher (has museum ID) | ✔(or) | ✔(or) | ||||
| Public voucher (agreed institution) | ✔(or) | ✔(or) | ||||
| Public voucher (agreed voucher type) | ✔(or) | ✔(or) | ||||
| Collection (Date) | ✔ | |||||
| Collection (Site) | ✔(or) | |||||
| Collection (GPS coordinate) | ✔(or) | |||||
| Collection (Sector) | ✔(or) | |||||
| Collection (Region) | ✔(or) | |||||
| Collector named | ✔ | |||||
While it is technically possible to conduct manual curation using common spreadsheets, this approach becomes impractical and error-prone for large and metadata-rich datasets (
To support this critical second phase, we developed a dedicated, browser-based curation tool tailored to the needs of taxonomic experts.
Here we introduce the design and functionality of the Library Curation Tool, and highlight its potential role in producing high-quality barcode reference data for DNA-based species identification. In doing so, we aim to provide a scalable, transparent, and expert-driven solution for curating large and complex barcode datasets, particularly in the context of biodiversity research and monitoring initiatives.
BGE Library Curation Tool
The Libary Curation Tool allows curators to review, validate, and annotate BOLD-derived records using a structured, user-friendly interface (cf. Fig.
The manual curation process using this tool generally follows these steps:
Load dataset: Select the .db file containing the data for the target taxonomic group generated by the pre-curation pipeline.
Filter and search: Narrow down the dataset by species, BIN, or other metadata fields.
Review records: Examine species names, BIN assignments, metadata quality, and potential conflicts.
Assign status: Mark each record as valid, invalid, or excluded; correct species names where needed.
Add notes: Document curation decisions with curator comments.
Monitor statistics: Use dynamic counters to track BIN-sharing/splitting events and curation completeness.
Export results: Save curated data as .csv along with a changes.log file for audit purposes.
Biodiversity Genomics Europe (Grant no.101059492) is funded by Horizon Europe (
The Library Curation Tool (cf. Fig.
HTML table of the Library Curation Tool with taxonomic records from BOLD and additional curation specific metadata (like: url, Ranking, country_representative, BAGS, Status, Reason Name Correction, Correct Species Name, Curator Notes). Records with grey background are pre-selected for reference library and need no action by the user for getting them added to the reference library. However, users can validate non-pre-selected records, which will get a green background or invalidate pre-selected records, which will have a red background.
All actions performed in the Library Curation Tool are documented in log files including timestamps, the executed action, and all information provided by the experts. These log files are analysed using a dedicated script (https://github.com/FabianDeister/BGE_library_curation_tool_log_processing) together with the output of the automated pipeline, whereby timestamps ensure that only the most recent version of each change is retained. In this way, both automatically pre-validated and manually reviewed records are merged to form the curated reference library. The resulting curated datasets provide the basis for the development of curated European DNA barcode reference libraries within the Biodiversity Genomics Europe project. In the longer term, the project aims to make curated outputs and expert annotations accessible through shared infrastructures, including BOLD and BOLD-Europe, allowing curation decisions to contribute to community-maintained reference resources. The exact mechanisms for integration and publication are currently under discussion with BOLD partners.
The application is implemented using standard web technologies and can be run entirely on a local computer, without internet access (aside from optional loading of remote CSS assets). It includes the following components:
Backend: The backend logic is handled by a lightweight Node.js (
Frontend: The user interface is built in index.html using HTML, CSS, and JavaScript, and runs in a modern web browser (e.g. Chrome, Firefox). It utilizes the DataTables library to provide interactive, searchable, and paginated tables. Custom JavaScript code supports advanced functionalities such as row coloring based on status, BIN visualization, in-table dropdowns for status selection, curator note entry, and per-record submission.
Database Input: The tool operates on local .db database files. These files are placed in the data/ subdirectory and loaded dynamically through a dataset selector. Each file corresponds to a taxonomic group and contains hundreds to thousands of records to be curated.
Execution Environment: The application is platform-independent and distributed as a self-contained folder Fig.
Export and Audit: All curation actions (status changes, species name updates, notes) are logged in a changes.log file, ensuring transparency and reproducibility. Curators can export their results (for own purposes) in .csv format and submit the log file as a standardized feedback mechanism. The tool prevents structural errors common in spreadsheet-based curation by enforcing consistent fields and controlled input types.
Dynamic Scoring: The tool includes dynamic logic for recalculating BAGS scores and BIN statistics in real time. This allows experts to see how their actions (e.g. excluding a species or marking a record invalid) influence BIN-sharing, BIN-splitting, and representative selection.
In summary, the Library Curation Tool is a locally hosted, browser-accessible interface purpose-built for scalable expert curation of DNA barcode data. It bridges the gap between automated pipeline output and final expert-reviewed reference libraries, facilitating the creation of high-quality, FAIR-compliant resources for molecular biodiversity research.
There are several levels of support for the user. First, there is a user manual within the main folder of the curation tool. Second, contextual help is provided the user interface by red question marks - clicking on them opens a menu with additional information. Third, a video tutorial and FAQ section are available on the project website (https://bge-barcoding.github.io/manual-curation/). An overview of the associated github repositories is presented in Table
| url / link | Description |
| Curation Tool - Linux and macOS version | |
|
https://github.com/bge-barcoding/BGE_library_curation_tool_win |
Curation Tool - Windows version |
| BOLD Library Curation Pipeline | |
| iBOL Europe BOLD Curation Datasets | |
| https://github.com/FabianDeister/BGE_library_curation_tool_log_processing | Curation Tool - Log Processing |
Although the current implementation operates locally on the curator's computer, the workflow is designed around standardized data structures, controlled vocabularies, reproducible log files, and version-controlled outputs. These features facilitate the transparent exchange of curation decisions and support future integration into shared reference data infrastructures. Rather than promoting isolated local reference databases, the long-term objective is to enable expert contributions from distributed specialists to be consolidated into community-curated, FAIR-compliant reference libraries that are accessible and reusable across projects, institutions, and countries.
Taxonomic experts will curate records from BOLD Systems that have been pre-curated using this pipeline: https://github.com/bge-barcoding/bold-library-curation.