- Home
- Search Results
- Page 1 of 1
Search for: All records
-
Total Resources2
- Resource Type
-
0000000002000000
- More
- Availability
-
20
- Author / Contributor
- Filter by Author / Creator
-
-
Allen, Julie M (1)
-
Baxter, David G. (1)
-
Blum, Stanley D. (1)
-
Bolmgren, Kjell (1)
-
Carter, J. Richard (1)
-
Davis, Charles C. (1)
-
Dean, Ellen (1)
-
Denny, Ellen G. (1)
-
Denslow, Michael W (1)
-
Denslow, Michael W. (1)
-
Ellwood, Elizabeth R. (1)
-
Gallinat, Amanda S. (1)
-
Gilbert, Ed (1)
-
Guralnick, Robert (1)
-
Guralnick, Robert P (1)
-
Haston, Elspeth M. (1)
-
LaFrance, Raphael (1)
-
Mazer, Susan J. (1)
-
Mishler, Brent D. (1)
-
Morris, Ashley B. (1)
-
- Filter by Editor
-
-
& Spizer, S. M. (0)
-
& . Spizer, S. (0)
-
& Ahn, J. (0)
-
& Bateiha, S. (0)
-
& Bosch, N. (0)
-
& Brennan K. (0)
-
& Brennan, K. (0)
-
& Chen, B. (0)
-
& Chen, Bodong (0)
-
& Drown, S. (0)
-
& Ferretti, F. (0)
-
& Higgins, A. (0)
-
& J. Peters (0)
-
& Kali, Y. (0)
-
& Ruiz-Arias, P.M. (0)
-
& S. Spitzer (0)
-
& Sahin. I. (0)
-
& Spitzer, S. (0)
-
& Spitzer, S.M. (0)
-
(submitted - in Review for IEEE ICASSP-2024) (0)
-
-
Have feedback or suggestions for a way to improve these results?
!
Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Abstract PremiseOne of the slowest steps in digitizing natural history collections is converting labels associated with specimens into a digital data record usable for collections management and research. Here, we address how herbarium specimen labels can be converted into digital data records via extraction into standardized Darwin Core fields. MethodsWe first showcase the development of a rule‐based approach and compare outcomes with a large language model–based approach, in particular ChatGPT4. We next quantified omission and commission error rates across target fields for a set of labels transcribed using optical character recognition (OCR) for both approaches. For example, we find that ChatGPT4 often creates field names that are not Darwin Core compliant while rule‐based approaches often have high commission error rates. ResultsOur results suggest that these approaches each have different strengths and limitations. We therefore developed an ensemble approach that leverages the strengths of each individual method and documented that ensembling strongly reduced overall information extraction errors. DiscussionThis work shows that an ensemble approach has particular value for creating high‐quality digital data records, even for complicated label content. While human validation is still needed to ensure the best possible quality, automated approaches can speed digitization of herbarium specimen labels and are likely to be broadly usable for all natural history collection types.more » « less
-
Yost, Jennifer M.; Sweeney, Patrick W.; Gilbert, Ed; Nelson, Gil; Guralnick, Robert; Gallinat, Amanda S.; Ellwood, Elizabeth R.; Rossington, Natalie; Willis, Charles G.; Blum, Stanley D.; et al (, Applications in Plant Sciences)
An official website of the United States government
