skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: Connecting the Dots: Aligning human capacity through networks toward a globally interoperable Digital Extended Specimen (DES) infrastructure
Thanks to substantial support for biodiversity data mobilization in recent decades, billions of occurrence records are openly available, documenting life on Earth and enabling timely research, awareness raising, and policy-making. Initiatives across local to global scales have been separately funded to serve different, yet often overlapping audiences of data users, and have developed a variety of platforms and infrastructures to meet the needs of these audiences. The independent progress of biodiversity data providers has led to innovations as well as challenges for the community at large as we move towards connecting and linking a diversity of information from disparate sources as Digital Extended Specimens (DES). Recognizing a need for deeper and more frequent opportunities for communication and collaboration across the globe, an ad-hoc group of representatives of various international, national, and regional organizations have been meeting virtually since 2020 to provide a forum for updates, announcements, and shared progress. This group is provisionally named International Partners for the Digital Extended Specimen (IPDES), and is guided by these four concepts: Biodiversity, Connection, Knowledge and Agency. Participants in IPDES include representatives of the Global Biodiversity Information Facility (GBIF), Integrated Digitized Biocollections (iDigBio), American Institute of Biological Sciences (AIBS), Biodiversity Collections Network (BCoN), Natural Science Collections Alliance (NSCA), Distributed System of Scientific Collections (DiSSCo), Atlas of Living Australia (ALA), Biodiversity Information Standards (TDWG), Society for the Preservation of Natural History Collections (SPNHC), National Specimen Information Infrastructure of China (NSII), and South African National Biodiversity Institute (SANBI), as well as individuals involved with biodiversity informatics initiatives, natural science collections, museums, herbaria, and universities. Our global partners group strives to increase representation from around the globe as we aim to enable research that contributes to novel discoveries and addresses the societal challenges leading to the biodiversity crisis. Our overarching mission is to expand on the community-driven successes to connect biodiversity data and knowledge through coordination of a globally integrated network of stakeholders to enable an extensible technical and social infrastructure of data, tools, and working practices in support of our vision. The main work of our group thus far includes publishing a paper on the Digital Extended Specimen (Hardisty et al. 2022), organizing and hosting an array of activities at conferences, and asynchronous online work and forum-based exchanges. We aim to advance discussion on topics of broad interest to our community such as social and technical capacity building, broadening participation, expanding social and data networks, improving data models and building a backbone for the DES, and identifying international funding solutions. This presentation will highlight some of these activities and detail progress towards a roadmap for the development of the human network and technical infrastructure necessary to support the DES. It provides an opportunity for feedback from and engagement by stakeholder communities such as TDWG and other initiatives with a focus on data standards and biodiversity informatics, as we solidify our plans for the future in support of integrated and interconnected biodiversity data and credit for those doing the work.  more » « less
Award ID(s):
2148939
PAR ID:
10504184
Author(s) / Creator(s):
; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; more » ; ; ; ; ; « less
Publisher / Repository:
Biodiversity Information Science and Standards
Date Published:
Journal Name:
Biodiversity Information Science and Standards
Volume:
7
ISSN:
2535-0897
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. International collaboration between collections, aggregators, and researchers within the biodiversity community and beyond is becoming increasingly important in our efforts to support biodiversity, conservation and the life of the planet. The social, technical, logistical and financial aspects of an equitable biodiversity data landscape – from workforce training and mobilization of linked specimen data, to data integration, use and publication – must be considered globally and within the context of a growing biodiversity crisis. In recent years, several initiatives have outlined paths forward that describe how digital versions of natural history specimens can be extended and linked with associated data. In the United States, Webster (2017) presented the “extended specimen”, which was expanded upon by Lendemer et al. (2019) through the work of the Biodiversity Collections Network (BCoN). At the same time, a “digital specimen” concept was developed by DiSSCo in Europe (Hardisty 2020). Both the extended and digital specimen concepts depict a digital proxy of an analog natural history specimen, whose digital nature provides greater capabilities such as being machine-processable, linkages with associated data, globally accessible information-rich biodiversity data, improved tracking, attribution and annotation, additional opportunities for data use and cross-disciplinary collaborations forming the basis for FAIR (Findable, Accessible, Interoperable, Reproducible) and equitable sharing of benefits worldwide, and innumerable other advantages, with slight variation in how an extended or digital specimen model would be executed. Recognizing the need to align the two closely-related concepts, and to provide a place for open discussion around various topics of the Digital Extended Specimen (DES; the current working name for the joined concepts), we initiated a virtual consultation on the discourse platform hosted by the Alliance for Biodiversity Knowledge through GBIF. This platform provided a forum for threaded discussions around topics related and relevant to the DES. The goals of the consultation align with the goals of the Alliance for Biodiversity Knowledge: expand participation in the process, build support for further collaboration, identify use cases, identify significant challenges and obstacles, and develop a comprehensive roadmap towards achieving the vision for a global specification for data integration. In early 2021, Phase 1 launched with five topics: Making FAIR data for specimens accessible; Extending, enriching and integrating data; Annotating specimens and other data; Data attribution; and Analyzing/mining specimen data for novel applications. This round of full discussion was productive and engaged dozens of contributors, with hundreds of posts and thousands of views. During Phase 1, several deeper, more technical, or additional topics of relevance were identified and formed the foundation for Phase 2 which began in May 2021 with the following topics: Robust access points and data infrastructure alignment; Persistent identifier (PID) scheme(s); Meeting legal/regulatory, ethical and sensitive data obligations; Workforce capacity development and inclusivity; Transactional mechanisms and provenance; and Partnerships to collaborate more effectively. In Phase 2 fruitful progress was made towards solutions to some of these complex functional and technical long-term goals. Simultaneously, our commitment to open participation was reinforced, through increased efforts to involve new voices from allied and complementary fields. Among a wealth of ideas expressed, the community highlighted the need for unambiguous persistent identifiers and a dedicated agent to assign them, support for a fully linked system that includes robust publishing mechanisms, strong support for social structures that build trustworthiness of the system, appropriate attribution of legacy and new work, a system that is inclusive, removed from colonial practices, and supportive of creative use of biodiversity data, building a truly global data infrastructure, balancing open access with legal obligations and ethical responsibilities, and the partnerships necessary for success. These two consultation periods, and the myriad activities surrounding the online discussion, produced a wide variety of perspectives, strategies, and approaches to converging the digital and extended specimen concepts, and progressing plans for the DES -- steps necessary to improve access to research-ready data to advance our understanding of the diversity and distribution of life. Discussions continue and we hope to include your contributions to the DES in future implementation plans. 
    more » « less
  2. Abstract Natural history collections are repositories of biodiversity specimens that provide critical infrastructure for studies of mammals. Over the past 3 decades, digitization of collections has opened up the temporal and spatial properties of specimens, stimulating new data sharing, use, and training across the biodiversity sciences. These digital records are the cornerstones of an “extended specimen network,” in which the diverse data derived from specimens become digital, linked, and openly accessible for science and policy. However, still missing from most digital occurrences of mammals are their morphological, reproductive, and life-history traits. Unlocking this information will advance mammalogy, establish richer faunal baselines in an era of rapid environmental change, and contextualize other types of specimen-derived information toward new knowledge and discovery. Here, we present the Ranges Digitization Network (Ranges), a community effort to digitize specimen-level traits from all terrestrial mammals of western North America, append them to digital records, publish them openly in community repositories, and make them interoperable with complimentary data streams. Ranges is a consortium of 23 institutions with an initial focus on non-marine mammal species (both native and introduced) occurring in western Canada, the western United States, and Mexico. The project will establish trait data standards and informatics workflows that can be extended to other regions, taxa, and traits. Reconnecting mammalogists, museum professionals, and researchers for a new era of collections digitization will catalyze advances in mammalogy and create a community-curated trait resource for training and engagement with global conservation initiatives. 
    more » « less
  3. Abstract The early twenty-first century has witnessed massive expansions in availability and accessibility of digital data in virtually all domains of the biodiversity sciences. Led by an array of asynchronous digitization activities spanning ecological, environmental, climatological, and biological collections data, these initiatives have resulted in a plethora of mostly disconnected and siloed data, leaving to researchers the tedious and time-consuming manual task of finding and connecting them in usable ways, integrating them into coherent data sets, and making them interoperable. The focus to date has been on elevating analog and physical records to digital replicas in local databases prior to elevating them to ever-growing aggregations of essentially disconnected discipline-specific information. In the present article, we propose a new interconnected network of digital objects on the Internet—the Digital Extended Specimen (DES) network—that transcends existing aggregator technology, augments the DES with third-party data through machine algorithms, and provides a platform for more efficient research and robust interdisciplinary discovery. 
    more » « less
  4. Leal, JH; Bieler, R (Ed.)
    Among biocollections, mollusks are a particularly powerful resource for a wide range of studies, including biogeography, conservation, ecology, environmental monitoring, evolutionary biology, and systematics. U.S. mollusk collections are housed in stand-alone natural history museums, at universities, and in a variety of governmental and non-governmental institutions. Differing in their histories, specializations, and uses, they share common needs for long-term development, and collectively contribute to biodiversity knowledge at regional, national, and global scales. Commitment by dedicated staff, collectors, and volunteers, institutional investments, philanthropy, and governmental funding have built and maintained these collections and their support infrastructure. Efforts by the North American malacological collection community since the early 1970s led to coordination in database design but left the data isolated in individual institutions. Collection digitization developed through a combination of individual/institutional initiatives and federally supported projects funded by the National Science Foundation (NSF) and the Institute of Museum and Library Services (IMLS). Advances in digital technology enabled the shift toward nationally and globally unified collections. Networking and collaboration were greatly accelerated by NSF’s Advancing Digitization of Biodiversity Collections (ADBC) program, which created a central coordinating organization (iDigBio) and funded Thematic Collections Network (TCN) projects. One such TCN was developed to mobilize nearly 90% of the known U.S. museum-collections-based data of the U.S. Atlantic and Gulf coasts (Mobilizing Millions of Marine Mollusks of the Eastern Seaboard—ESB). The project, involving 16 museum collections (plus the Smithsonian Institution as federal partner), combines data from approximately 4.5 million specimens collected from the ESB region and makes them available to the TCN portal InvertEBase and other aggregators such as iDigBio and GBIF. In addition to fostering community and expanding the corpus of available digitized mollusk records through new data entry and georeferencing (GEOLocate, CoGe) and standardizing taxonomy, the project drove key innovations for the invertebrate collections community. For instance, it worked with the Biodiversity Information Standards (TDWG) group to create a new Darwin Core standard term, “Vitality”, expanded GEOLocate to support complex geospatial types, integrated global elevation and bathymetric datasets directly into georeferencing workflow, and developed various education and outreach public outreach products. Synthesizing from the 15 following articles with individual histories of ESB-participating mollusk collections, several topics are discussed—such as what defines a “good” mollusk collection in the digital age and the importance of federal support for this national resource. 
    more » « less
  5. Among biocollections, mollusks are a particularly powerful resource for a wide range of studies, including biogeography, conservation, ecology, environmental monitoring, evolutionary biology, and systematics. U.S. mollusk collections are housed in stand-alone natural history museums, at universities, and in a variety of governmental and non-governmental institutions. Differing in their histories, specializations, and uses, they share common needs for long-term development, and collectively contribute to biodiversity knowledge at regional, national, and global scales. Commitment by dedicated staff, collectors, and volunteers, institutional investments, philanthropy, and governmental funding have built and maintained these collections and their support infrastructure. Efforts by the North American malacological collection community since the early 1970s led to coordination in database design but left the data isolated in individual institutions. Collection digitization developed through a combination of individual/institutional initiatives and federally supported projects funded by the National Science Foundation (NSF) and the Institute of Museum and Library Services (IMLS). Advances in digital technology enabled the shift toward nationally and globally unified collections. Networking and collaboration were greatly accelerated by NSF’s Advancing Digitization of Biodiversity Collections (ADBC) program, which created a central coordinating organization (iDigBio) and funded Thematic Collections Network (TCN) projects. One such TCN was developed to mobilize nearly 90% of the known U.S. museum-collections-based data of the U.S. Atlantic and Gulf coasts (Mobilizing Millions of Marine Mollusks of the Eastern Seaboard—ESB). The project, involving 16 museum collections (plus the Smithsonian Institution as federal partner), combines data from approximately 4.5 million specimens collected from the ESB region and makes them available to the TCN portal InvertEBase and other aggregators such as iDigBio and GBIF. In addition to fostering community and expanding the corpus of available digitized mollusk records through new data entry and georeferencing (GEOLocate, CoGe) and standardizing taxonomy, the project drove key innovations for the invertebrate collections community. For instance, it worked with the Biodiversity Information Standards (TDWG) group to create a new Darwin Core standard term, “Vitality”, expanded GEOLocate to support complex geospatial types, integrated global elevation and bathymetric datasets directly into georeferencing workflow, and developed various education and outreach public outreach products. Synthesizing from the 15 following articles with individual histories of ESB-participating mollusk collections, several topics are discussed—such as what defines a “good” mollusk collection in the digital age and the importance of federal support for this national resource. 
    more » « less