G. R.A.F. and R.T.L. wrote the manuscript; all authors edited the manuscript. All authors approved the final version of the manuscript. Open Research Badges This article has earned an Open Data badge for making publicly available the digitally shareable data necessary to reproduce the reported results. The data are available at https://doi.org/10.5281/zenodo.8336468 and https://doi.org/10.5281/zenodo.8349315 . Supporting information Appendix S1 . Guide to generalizing natural language processing and overcoming challenges. Click here for additional data file. ACKNOWLEDGMENTS We first thank the collective efforts of taxonomists over hundreds of years; the automated approach described here fo
Open resource ↗10.5281/zenodo.8336468 · lines:98-129Unverified paper record
FloraTraiter: Automated parsing of traits from descriptive biodiversity literature.
Applications in plant sciences · 18 Jan 2024 · 10.1002/aps3.11563
Abstract
Premise Plant trait data are essential for quantifying biodiversity and function across Earth, but these data are challenging to acquire for large studies. Diverse strategies are needed, including the liberation of heritage data locked within specialist literature such as floras and taxonomic monographs. Here we report FloraTraiter, a novel approach using rule-based natural language processing (NLP) to parse computable trait data from biodiversity literature. Methods FloraTraiter was implemented through collaborative work between programmers and botanical experts and customized for both online floras and scanned literature. We report a strategy spanning optical character recognition, recognition of taxa, iterative building of traits, and establishing linkages among all of these, as well as curational tools and code for turning these results into standard morphological matrices. Results Over 95% of treatment content was successfully parsed for traits with Conclusions We identify strategies, applications, tips, and challenges that we hope will facilitate future similar efforts to produce large open-source trait data sets for broad community reuse. Largely automated tools like FloraTraiter will be an important addition to the toolkit for assembling trait data at scale.
Plant phenotyping relevance
植物の形態形質を文献から自動抽出するNLPツールの開発が中心であり、植物フェノタイピング手法として適格です。
abstractHere we report FloraTraiter, a novel approach using rule-based natural language processing (NLP) to parse computable trait data from biodiversity literature.
abstractWe report a strategy spanning optical character recognition, recognition of taxa, iterative building of traits, and establishing linkages among all of these, as well as curational tools and code for turning these results into standard morphological matrices.
Code and data availability
The paper's trait-extraction codebase (FloraTraiter) and the Fagales worked-example repository with extracted trait data are both publicly available on GitHub with archived Zenodo releases, as stated in the Data Availability Statement.
all authors edited the manuscript. All authors approved the final version of the manuscript. Open Research Badges This article has earned an Open Data badge for making publicly available the digitally shareable data necessary to reproduce the reported results. The data are available at https://doi.org/10.5281/zenodo.8336468 and https://doi.org/10.5281/zenodo.8349315 . Supporting information Appendix S1 . Guide to generalizing natural language processing and overcoming challenges. Click here for additional data file. ACKNOWLEDGMENTS We first thank the collective efforts of taxonomists over hundreds of years; the automated approach described here focuses on recent works, but this still subst
Open resource ↗10.5281/zenodo.8349315 · lines:98-129This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.