skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: Open Source Tools for Scaling Data Curation at QDR
This paper describes the development of services and tools for scaling data curation services at the Qualitative Data Repository (QDR). Through a set of open-source tools, semi-automated workflows, and extensions to the Dataverse platform, our team has built services for curators to efficiently and effectively publish collections of qualitatively derived data. The contributions we seek to make in this paper are as follows: 1. We describe ‘human-in-the-loop’ curation and the tools that facilitate this model at QDR; 2. We provide an in-depth discussion of the design and implementation of these tools, including applications specific to the Dataverse software repository, as well as standalone archiving tools written in R; and 3. We highlight the role of providing a service layer for data discovery and accessibility of qualitative data.  more » « less
Award ID(s):
1823950
PAR ID:
10219835
Author(s) / Creator(s):
; ;
Date Published:
Journal Name:
The code4lib journal
Issue:
49
ISSN:
1940-5758
Page Range / eLocation ID:
https://journal.code4lib.org/articles/15436
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Accessibility of research data to disabled users has received scant attention in literature and practice. In this paper we briefly survey the current state of accessibility for research data and suggest some first steps that repositories should take to make their holdings more accessible. We then describe in depth how those steps were implemented at the Qualitative Data Repository (QDR), a domain repository for qualitative social-science data. The paper discusses accessibility testing and improvements on the repository and its underlying software, changes to the curation process to improve accessibility, as well as efforts to retroactively improve the accessibility of existing collections. We conclude by describing key lessons learned during this process as well as next steps. 
    more » « less
  2. In this short practice paper, we introduce the public version of the Qualitative Data Repository’s (QDR) Curation Handbook. The Handbook documents and structures curation practices at QDR. We describe the background and genesis of the Handbook and highlight some of its key content. 
    more » « less
  3. Data sharing is increasingly an expectation in health research as part of a general move toward more open sciences. In the United States, in particular, the implementation of the 2023 National Institutes of Health Data Management and Sharing Policy has made it clear that qualitative studies are not exempt from this data sharing requirement. Recognizing this trend, the Palliative Care Research Cooperative Group (PCRC) realized the value of creating a de-identified qualitative data repository to complement its existing de-identified quantitative data repository. The PCRC Data Informatics and Statistics Core leadership partnered with the Qualitative Data Repository (QDR) to establish the first serious illness and palliative care qualitative data repository in the U.S. We describe the processes used to develop this repository, called the PCRC-QDR, as well as our outreach and education among the palliative care researcher community, which led to the first ten projects to share the data in the new repository. Specifically, we discuss how we co-designed the PCRC-QDR and created tailored guidelines for depositing and sharing qualitative data depending on the original research context, establishing uniform expectations for key components of relevant documentation, and the use of suitable access controls for sensitive data. We also describe how PCRC was able to leverage its existing community to recruit and guide early depositors and outline lessons learned in evaluating the experience. This work advances the establishment of best practices in qualitative data sharing. 
    more » « less
  4. How can authors using many individual pieces of qualitative data throughout a publication make their research transparent? In this paper we introduce Annotation for Transparent Inquiry (ATI), an approach to enhance transparency in qualitative research. ATI allows authors to connect specific passages in their publication with an annotation. These annotations provide additional information relevant to the passage and, when possible, include a link to one or more data sources underlying a claim; data sources are housed in a repository. After describing ATI’s conceptual and technological implementation, we report on its evaluation through a series of workshops conducted by the Qualitative Data Repository (QDR) and present initial results of the evaluation. The article ends with an outlook on next steps for the project. 
    more » « less
  5. Abstract Fueled by the explosion of (meta)genomic data, genome mining of specialized metabolites has become a major technology for drug discovery and studying microbiome ecology. In these efforts, computational tools like antiSMASH have played a central role through the analysis of Biosynthetic Gene Clusters (BGCs). Thousands of candidate BGCs from microbial genomes have been identified and stored in public databases. Interpreting the function and novelty of these predicted BGCs requires comparison with a well-documented set of BGCs of known function. The MIBiG (Minimum Information about a Biosynthetic Gene Cluster) Data Standard and Repository was established in 2015 to enable curation and storage of known BGCs. Here, we present MIBiG 2.0, which encompasses major updates to the schema, the data, and the online repository itself. Over the past five years, 851 new BGCs have been added. Additionally, we performed extensive manual data curation of all entries to improve the annotation quality of our repository. We also redesigned the data schema to ensure the compliance of future annotations. Finally, we improved the user experience by adding new features such as query searches and a statistics page, and enabled direct link-outs to chemical structure databases. The repository is accessible online at https://mibig.secondarymetabolites.org/. 
    more » « less