Abstract Driven by the big data science, material informatics has attracted enormous research interests recently along with many recognized achievements. To acquire knowledge of materials by previous experience, both feature descriptors and databases are essential for training machine learning (ML) models with high accuracy. In this regard, the electronic charge density ρ ( r ), which in principle determines the properties of materials at their ground state, can be considered as one of the most appropriate descriptors. However, the systematic electronic charge density ρ ( r ) database of inorganic materials is still in its infancy due to the difficulties in collecting raw data in experiment and the expensive first-principles based computational cost in theory. Herein, a real space electronic charge density ρ ( r ) database of 17,418 cubic inorganic materials is constructed by performing high-throughput density functional theory calculations. The displayed ρ ( r ) patterns show good agreements with those reported in previous studies, which validates our computations. Further statistical analysis reveals that it possesses abundant and diverse data, which could accelerate ρ ( r ) related machine learning studies. Moreover, the electronic charge density database will also assists chemical bonding identifications and promotes new crystal discovery in experiments.
more »
« less
The LISTEN principles for genetic sequence data governance and database engineering
Several international legal agreements include an ‘access and benefit-sharing’ (ABS) mechanism that attaches obligations to the use of genetic sequence data. These agreements are frequently subject to critique on the grounds that ABS is either fundamentally incompatible with the principles of open science, or technically challenging to implement in open scientific databases. Here, we argue that these critiques arise from a misinterpretation of the principles of open science and that both considerations can be addressed by a set of simple principles that link database engineering and governance. We introduce a checklist of six database design considerations, LISTEN: licensed, identified, supervised, transparent, enforced and non-exclusive, which can be readily adopted by both new and existing platforms participating in ABS systems. We also highlight how these principles can act in concert with familiar principles of open science, such as findable, accessible, interoperable and reusable (FAIR) data sharing.
more »
« less
- PAR ID:
- 10694257
- Publisher / Repository:
- Nature Genetics
- Date Published:
- Journal Name:
- Nature Genetics
- Volume:
- 57
- Issue:
- 9
- ISSN:
- 1061-4036
- Page Range / eLocation ID:
- 2099 to 2105
- Subject(s) / Keyword(s):
- databases genetic sequences pathogens Pandemic Agreement international law data sharing
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Involving the public in scientific discovery offers opportunities for engagement, learning, participation, and action. Since its launch in 2007, the CitSci.org platform has supported hundreds of community-driven citizen science projects involving thousands of participants who have generated close to a million scientific measurements around the world. Members using CitSci.org follow their curiosities and concerns to develop, lead, or simply participate in research projects. While professional scientists are trained to make ethical determinations related to the collection of, access to, and use of information, citizen scientists and practitioners may be less aware of such issues and more likely to become involved in ethical dilemmas. In this era of big and open data, where data sharing is encouraged and open science is promoted, privacy and openness considerations can often be overlooked. Platforms that support the collection, use, and sharing of data and personal information need to consider their responsibility to protect the rights to and ownership of data, the provision of protection options for data and members, and at the same time provide options for openness. This requires critically considering both intended and unintended consequences of the use of platforms, data, and volunteer information. Here, we use our journey developing CitSci.org to argue that incorporating customization into platforms through flexible design options for project managers shifts the decision-making from top-down to bottom-up and allows project design to be more responsive to goals. To protect both people and data, we developed—and continue to improve—options that support various levels of “open” and “closed” access permissions for data and membership participation. These options support diverse governance styles that are responsive to data uses, traditional and indigenous knowledge sensitivities, intellectual property rights, personally identifiable information concerns, volunteer preferences, and sensitive data protections. We present a typology for citizen science openness choices, their ethical considerations, and strategies that we are actively putting into practice to expand privacy options and governance models based on the unique needs of individual projects using our platform.more » « less
-
Data sharing and transparency are becoming more common across the social sciences. In this article, we provide an overview of ethical, methodological, and technological considerations and challenges when developing large video-based datasets intended to be shared across researchers. We cover data security, storage, and access as well as data documentation, tagging, and transcription. Our discussions are framed by our own efforts to create a secure and user-friendly database for the New Jersey Families Study, a two-week, in-home video study of 21 families with a 2- to 4-year-old child. In collecting over 11,470 hours of video data, the New Jersey Families Study is one of the very few large-scale video projects in the field of sociology. This project has provided us with a unique opportunity to explore video data management and data sharing techniques, particularly in light of a host of cutting-edge developments in data science.more » « less
-
Abstract Trait-based approaches are revolutionizing our understanding of high-diversity ecosystems by providing insights into the principles underlying key ecological processes, such as community assembly, species distribution, resilience, and the relationship between biodiversity and ecosystem functioning. In 2016, the Coral Trait Database advanced coral reef science by centralizing trait information for stony corals (i.e., Subphylum Anthozoa, Class Hexacorallia, Order Scleractinia). However, the absence of trait data for soft corals, gorgonians, and sea pens (i.e., Class Octocorallia) limits our understanding of ecosystems where these organisms are significant members and play pivotal roles. To address this gap, we introduce the Octocoral Trait Database, a global, open-source database of curated trait data for octocorals. This database houses species- and individual-level data, complemented by contextual information that provides a relevant framework for analyses. The inaugural dataset, OctocoralTraits v2.2, contains over 97,500 global trait observations across 98 traits and over 3,500 species. The database aims to evolve into a steadily growing, community-led resource that advances future marine science, with a particular emphasis on coral reef research.more » « less
-
Pickett, Brett E.; Jurado, Kellie (Ed.)ABSTRACT Data that catalogue viral diversity on Earth have been fragmented across sources, disciplines, formats, and various degrees of open sharing, posing challenges for research on macroecology, evolution, and public health. Here, we solve this problem by establishing a dynamically maintained database of vertebrate-virus associations, called The Global Virome in One Network (VIRION). The VIRION database has been assembled through both reconciliation of static data sets and integration of dynamically updated databases. These data sources are all harmonized against one taxonomic backbone, including metadata on host and virus taxonomic validity and higher classification; additional metadata on sampling methodology and evidence strength are also available in a harmonized format. In total, the VIRION database is the largest open-source, open-access database of its kind, with roughly half a million unique records that include 9,521 resolved virus “species” (of which 1,661 are ICTV ratified), 3,692 resolved vertebrate host species, and 23,147 unique interactions between taxonomically valid organisms. Together, these data cover roughly a quarter of mammal diversity, a 10th of bird diversity, and ∼6% of the estimated total diversity of vertebrates, and a much larger proportion of their virome than any previous database. We show how these data can be used to test hypotheses about microbiology, ecology, and evolution and make suggestions for best practices that address the unique mix of evidence that coexists in these data. IMPORTANCE Animals and their viruses are connected by a sprawling, tangled network of species interactions. Data on the host-virus network are available from several sources, which use different naming conventions and often report metadata in different levels of detail. VIRION is a new database that combines several of these existing data sources, reconciles taxonomy to a single consistent backbone, and reports metadata in a format designed by and for virologists. Researchers can use VIRION to easily answer questions like “Can any fish viruses infect humans?” or “Which bats host coronaviruses?” or to build more advanced predictive models, making it an unprecedented step toward a full inventory of the global virome.more » « less
An official website of the United States government

