Abstract Reproducibility is a foundational tenet of science. As artificial intelligence (AI) becomes increasingly embedded across science, the need to accurately document the provenance, structure, and behavior of training data, models, and workflows grows correspondingly. Metadata, understood as explicit and structured knowledge about data and related entities, is a critical yet often underexamined component of AI systems that helps address this need. High‐quality metadata describing datasets, models, and workflows supports the FAIR (Findable, Accessible, Interoperable, Reusable) principles, strengthens reproducibility, and enables evaluation of AI‐readiness by making data and models interpretable, traceable, and structurally consistent. Despite its central importance, sustained discussion of metadata as a core component of AI infrastructure remains limited. This article examines the role of metadata in advancing AI‐enabled research, focusing on the design, implementation, and operationalization of metadata standards within an evolving metadata ecosystem. We first present a conceptual view of the metadata ecosystem, framed by data structure, data value, data encoding, and syntax standards, as a foundation for understanding how metadata enables AI and how AI contributes to metadata generation and refinement. We then introduce four case studies that illustrate how metadata can be generated, refined, and leveraged within AI workflows. The discussion synthesizes the cases, highlights limitations, including metadata quality challenges and the role of structured constraints in addressing AI errors, and relates each case to the metadata ecosystem dimensions it engages. Taken together, these cases illustrate that metadata is not a peripheral add‐on but an essential component of AI‐ready, FAIR‐aligned, transparent, and reproducible research systems.
more »
« less
AI-ready data in space science and solar physics: problems, mitigation and action plan
In the domain of space science, numerous ground-based and space-borne data of various phenomena have been accumulating rapidly, making analysis and scientific interpretation challenging. However, recent trends in the application of artificial intelligence (AI) have been shown to be promising in the extraction of information or knowledge discovery from these extensive data sets. Coincidentally, preparing these data for use as inputs to the AI algorithms, referred to as AI-readiness, is one of the outstanding challenges in leveraging AI in space science. Preparation of AI-ready data includes, among other aspects: 1) collection (accessing and downloading) of appropriate data representing the various physical parameters associated with the phenomena under study from different repositories; 2) addressing data formats such as conversion from one format to another, data gaps, quality flags and labeling; 3) standardizing metadata and keywords in accordance with NASA archive requirements or other defined standards; 4) processing of raw data such as data normalization, detrending, and data modeling; and 5) documentation of technical aspects such as processing steps, operational assumptions, uncertainties, and instrument profiles. Making all existing data AI-ready within a decade is impractical and data from future missions and investigations exacerbates this. This reveals the urgency to set the standards and start implementing them now. This article presents our perspective on the AI-readiness of space science data and mitigation strategies including definition of AI-readiness for AI applications; prioritization of data sets, storage, and accessibility; and identifying the responsible entity (agencies, private sector, or funded individuals) to undertake the task.
more »
« less
- Award ID(s):
- 2026579
- PAR ID:
- 10661255
- Publisher / Repository:
- Frontiers in Astronomy and Space Sciences
- Date Published:
- Journal Name:
- Frontiers in Astronomy and Space Sciences
- Volume:
- 10
- ISSN:
- 2296-987X
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
A team of literacy, science, and theatre educators have been working to engage children in an urban public school system in the United States through embodied performances, where students embody and dramatise science ideas. This study focuses on one fourth‐grade classroom when instruction was done remotely due to Covid‐19. Children in the class were asked to compose videos of themselves acting out and/or exploring science phenomena and concepts, and we analysed the affordances of these multimodal compositions. We situate the need for this study in claims from the Next Generation Science Standards that literacy skills are necessary to build and communicate science knowledge. In doing so, we center social semiotics perspectives that conceive of composition broadly as production‐oriented processes drawing from various semiotic resources. The multimodal compositions in Mr. M's science class included both primarily embodied compositions and primarily digital compositions, and we elaborate on one focal example of each in the findings. Intertwined affordances of the focal children and their classmates' multimodal science compositions include opportunities to creatively engage with and negotiate science ideas, to draw from personal and social knowledge during meaning‐making, and to intentionally make rhetorical choices.more » « less
-
Buildings account for nearly 40% of US energy consumption, making effective auditing and retrofitting of building envelopes essential to reducing energy use. However, traditional energy audits remain costly, labor-intensive, and difficult to scale across diverse buildings. This paper proposes a novel workflow that leverages multimodal AI to prepare structured building energy modeling (BEM) inputs directly from close-range RGB and infrared images. The workflow integrates three components: 1) close-range image feature fusion using CapsLab-based thermal anomaly segmentation and feature matching; 2) Neural Radiance Fields (NeRF) for rendering full-scale, photorealistic façades from fragmented image inputs; and 3) LLaVA prompt engineering, guided by vision-related variables screened from the U.S. Energy Information Administration’s 2018 Commercial Buildings Energy Consumption Survey (CBECS) codebook, to extract standardized enclosure properties in structured formats. A pilot study conducted on the D.M. Smith building at Georgia Tech campus demonstrates that the proposed pipeline can accurately identify attributes such as wall construction type, roof material, and number of stories while maintaining consistency with energy auditing standards. This study highlights the potential of AI-assisted workflows to automate key aspects of building audits, reduce labor costs, and generate reliable, simulation-ready data for tools such as EnergyPlus.more » « less
-
AI decision-support tools typically offer a fixed type of assistance, like AI recommendations and explanations, regardless of the specific decision, individual, or broader context. This fixed design has been shown to hinder both human-AI decision accuracy and human skill improvement in the task. We posit that AI assistance needs to be dynamic, changing in response to contextual factors (e.g., AI uncertainty, task difficulty), individual differences, and specified objectives (e.g., decision accuracy, skill improvement). To enable such adaptive support, we propose reinforcement learning (RL) as a general approach for modeling human-AI decision-making to optimize human-AI interaction for diverse objectives. RL enables optimizing various objectives in AI-assisted decision-making by tailoring and adaptively providing decision support to humans --- the right type of assistance, to the right person, at the right time. We instantiated our approach with two objectives: human-AI accuracy on the decision-making task and human skill improvement (i.e., learning about the task) and learned decision support policies from previous human-AI interaction data. We compared the optimized policies against several baselines in AI-assisted decision-making. Across two experiments (N = 316 and N = 964), our results consistently demonstrated that people interacting with policies optimized for accuracy achieve significantly higher accuracy --- and even human-AI complementarity --- compared to those interacting with any other type of AI support. Our results further indicated that human learning was more difficult to optimize than accuracy. While the policies learned the best available actions to optimize learning, participants who interacted with learning-optimized policies showed significant learning improvement only at times. Our research (1) demonstrates offline RL to be a promising approach to model the dynamics of human-AI decision-making, leading to policies that may optimize various objectives and provide novel insights about the AI-assisted decision-making space, and (2) emphasizes the importance of considering skill improvement and other human-centric objectives beyond accuracy in AI-assisted decision-making, opening up the novel research challenge of optimizing human-AI interaction for such objectives.more » « less
-
Understanding the world around us is a growing necessity for the whole public, as citizens are required to make informed decisions in their everyday lives about complex issues. Systems thinking (ST) is a promising approach for developing solutions to various problems that society faces and has been acknowledged as a crosscutting concept that should be integrated across educational science disciplines. However, studies show that engaging students in ST is challenging, especially concerning aspects like change over time and feedback. Using computational system models and a system dynamics approach can support students in overcoming these challenges when making sense of complex phenomena. In this paper, we describe an empirical study that examines how 10th grade students engage in aspects of ST through computational system modeling as part of a Next Generation Science Standards-aligned project-based learning unit on chemical kinetics. We show students’ increased capacity to explain the underlying mechanism of the phenomenon in terms of change over time that goes beyond linear causal relationships. However, student models and their accompanying explanations were limited in scope as students did not address feedback mechanisms as part of their modeling and explanations. In addition, we describe specific challenges students encountered when evaluating and revising models. In particular, we show epistemological barriers to fruitful use of real-world data for model revision. Our findings provide insights into the opportunities of a system dynamics approach and the challenges that remain in supporting students to make sense of complex phenomena and nonlinear mechanisms.more » « less
An official website of the United States government

