Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Brandt, Steven; Bradley, Shannon (Ed.)The SAGE Suite–SAGE1, SAGE2, and SAGE3–translates advances in visualization, cyberinfrastructure, and human-computer interaction into an open, scalable platform that aligns with embodied cognition to support collaborative, spatial reasoning on large displays and personal devices. Over two decades and hundreds of deployed walls worldwide, SAGE has enabled scientists, educators, and students to juxtapose heterogeneous media, sustain shared context, and accelerate sensemaking across the research lifecycle. This paper contributes: (1) a synthesis of the Suite’s translational impact across domains–from biology and atmospheric science to disaster management, health care, public outreach and workforce development; (2) a comparative framing of SAGE3 (the Smart Amplified Group Environment) among Computer Supported Cooperative Work and infinite-canvas tools; (3) the design rationale and user experience foundations of SAGE3’s “spatial thinking operating system,” including boards, rooms, wall viewports, and multi-user attention/flow mechanisms; (4) a modular architecture that delivers low-latency synchronization, extensibility via plugins, and privacy-aware deployment; and (5) a paradigm for human–Artificial Intelligence (AI) collaboration that spatializes notebooks and conversational workflows, enabling multi-user, multi-AI interaction grounded in shared visual context. We also surface systemic challenges in recognizing software-as-instrument within academic incentives and document emergent usage patterns spanning synchronous/asynchronous, co-located/distributed work. SAGE3 demonstrates how open, research-driven cyberinfrastructure can couple spatial cognition with collective intelligence to advance scientific collaboration and decision-making.more » « lessFree, publicly-accessible full text available March 1, 2027
-
Abstract BackgroundThe annotation of protein sequences in public databases has long posed a challenge in molecular biology. This issue is particularly acute for viral proteins, which demonstrate limited homology to known proteins when using alignment, k-mer, or profile-based homology search approaches. A novel methodology employing Large Language Models (LLMs) addresses this methodological challenge by annotating protein sequences based on embeddings. ResultsCentral to our contribution is the soft alignment algorithm, drawing from traditional protein alignment but leveraging embedding similarity at the amino acid level to bypass the need for conventional scoring matrices. This method not only surpasses pooled embedding-based models in efficiency but also in interpretability, enabling users to easily trace homologous amino acids and delve deeper into the alignments. Far from being a black box, our approach provides transparent, BLAST-like alignment visualizations, combining traditional biological research with AI advancements to elevate protein annotation through embedding-based analysis while ensuring interpretability. Tests using the Virus Orthologous Groups and ViralZone protein databases indicated that the novel soft alignment approach recognized and annotated sequences that both blastp and pooling-based methods, which are commonly used for sequence annotation, failed to detect. ConclusionThe embeddings approach shows the great potential of LLMs for enhancing protein sequence annotation, especially in viral genomics. These findings present a promising avenue for more efficient and accurate protein function inference in molecular biology.more » « less
-
With the emergence of Artificial Intelligence, it’s becoming essential for everyone—not just scientists and students—to harness its potential to stay competitive, think more critically, and drive innovation in a rapidly evolving world. SAGE3 is an open-source platform designed to help individuals and teams collaborate effectively—with each other and with AI—to accelerate the process of understanding, problem-solving, and discovery. It empowers everyday citizens to become smarter and more innovative by making complex information more accessible and actionable. Developed from over 20 years of National Science Foundation–funded research, SAGE3 is grounded in a deep understanding of how people work together across disciplines and interact with diverse streams of data. SAGE3 supports translational and convergent research, making it ideal for integrating insights from science, technology, community knowledge, and policy to tackle real-world challenges. It enables people to work with large and varied information sources—collaborating seamlessly with AI to reach decisions more quickly, clearly, and confidently. Whether working side-by-side on expansive shared display walls or contributing remotely from a laptop—at home, at work, or while traveling—SAGE3 enables flexible, co-located and distributed collaboration. It transforms static data into shared understanding, powering more informed, creative, and collective decision-making for all.more » « less
-
Current computational notebooks, such as Jupyter, are a popular tool for data science and analysis. However, they use a 1D list structure for cells that introduces and exacerbates user issues, such as messiness, tedious navigation, inefficient use of large screen space, performance of non-linear analyses, and presentation of non-linear narratives. To ameliorate these issues, we designed a prototype extension for Jupyter Notebooks that enables 2D organization of computational notebook cells into multiple columns. In this paper, we present two evaluative studies to determine whether such “2D computational notebooks” provide advantages over the current computational notebook structure. From these studies, we found empirical evidence that our multi-olumn 2D computational notebooks provide enhanced efficiency and usability. We also gathered design feedback which may inform future works. Overall, the prototype was positively received, with some users expressing a clear preference for 2D computational notebooks even at this early stage of development.more » « less
-
Current computational notebooks, such as Jupyter, are a popular tool for data science and analysis. However, they use a 1D list structure for cells that introduces and exacerbates user issues, such as messiness, tedious navigation, inefficient use of large screen space, performance of non-linear analyses, and presentation of non-linear narratives. To ameliorate these issues, we designed a prototype extension for Jupyter Notebooks that enables 2D organization of computational notebook cells into multiple columns. In this paper, we present two evaluative studies to determine whether such “2D computational notebooks” provide advantages over the current computational notebook structure. From these studies, we found empirical evidence that our multi-olumn 2D computational notebooks provide enhanced efficiency and usability. We also gathered design feedback which may inform future works. Overall, the prototype was positively received, with some users expressing a clear preference for 2D computational notebooks even at this early stage of development.more » « less
-
Current computational notebooks, such as Jupyter, are a popular tool for data science and analysis. However, they use a 1D list structure for cells that introduces and exacerbates user issues, such as messiness, tedious navigation, inefficient use of large screen space, performance of non-linear analyses, and presentation of non-linear narratives. To ameliorate these issues, we designed a prototype extension for Jupyter Notebooks that enables 2D organization of computational notebook cells into multiple columns. In this paper, we present two evaluative studies to determine whether such “2D computational notebooks” provide advantages over the current computational notebook structure. From these studies, we found empirical evidence that our multi-olumn 2D computational notebooks provide enhanced efficiency and usability. We also gathered design feedback which may inform future works. Overall, the prototype was positively received, with some users expressing a clear preference for 2D computational notebooks even at this early stage of development.more » « less
-
Current computational notebooks, such as Jupyter, are a popular tool for data science and analysis. However, they use a 1D list structure for cells that introduces and exacerbates user issues, such as messiness, tedious navigation, inefficient use of large screen space, performance of non-linear analyses, and presentation of non-linear narratives. To ameliorate these issues, we designed a prototype extension for Jupyter Notebooks that enables 2D organization of computational notebook cells into multiple columns. In this paper, we present two evaluative studies to determine whether such “2D computational notebooks” provide advantages over the current computational notebook structure. From these studies, we found empirical evidence that our multi-olumn 2D computational notebooks provide enhanced efficiency and usability. We also gathered design feedback which may inform future works. Overall, the prototype was positively received, with some users expressing a clear preference for 2D computational notebooks even at this early stage of development.more » « less
-
In collaboration with the Center for Microbiome Analysis through Island Knowledge and Investigations (C-MĀIKI), the Hawaii EPSCoR Ike Wai project and the Hawaii Data Science Institute, a new science gateway, the C-MĀIKI gateway, was developed to support modern, interoperable and scalable microbiome data analysis. This gateway provides a web-based interface for accessing high-performance computing resources and storage to enable and support reproducible microbiome data analysis. The C-MĀIKI gateway is accelerating the analysis of microbiome data for Hawaii through ease of use and centralized infrastructure.more » « less
-
Translational software research bridges the gap between scientific innovations and practical applications, driving impactful societal advancements. However, developing such software is challenging due to interdisciplinary collaboration, technology adoption, and postfunding sustainability. This article presents the experiences and insights of the Scalable Adaptive Graphics Environment (SAGE) team, which has spent two decades developing translational, cross-disciplinary, collaboration tools to benefit computational science research. With a focus on SAGE and its next-generation iterations, we explore the inherent challenges in translational research, such as fostering cross-disciplinary collaboration, motivating technology adoption, and ensuring postfunding product sustainability. We also discuss the roles of funding agencies, policymakers, and academic institutions in promoting translational research. Although the journey is fraught with challenges, the societal impact and satisfaction derived from translational research underscore its significance in the broader scientific landscape. This article aims to encourage further conversation and the development of effective models for translational software projects.more » « less
-
Protein language models (pLMs) have revolutionized computational biology by generating rich protein vector representations, or embeddings—enabling major advancements inde novoprotein design, structure prediction, variant effect analysis, and evolutionary studies. Despite these breakthroughs, current pLMs often exhibit biases against proteins from underrepresented species, with viral proteins being particularly affected, frequently referred to as the “dark matter” of the biological world due to their vast diversity and ubiquity, yet sparse representation in training datasets. Here, we show that fine-tuning pre-trained pLMs on viral protein sequences, using diverse learning frameworks and parameter-efficient strategies, significantly enhances representation quality and improves performance on downstream tasks. To support further research, we provide source code for fine-tuning pLMs and benchmarking embedding quality. By enabling more accurate modeling of viral proteins, our approach advances tools for understanding viral biology, combating emerging infectious diseases, and driving biotechnological innovation.more » « less
An official website of the United States government

Full Text Available