Scientific breakthroughs in biomolecular methods and improvements in hardware technology have shifted from a single long-running simulation to a large set of shorter simulations running simultaneously, called an ensemble. In an ensemble, each independent simulation is usually coupled with several analyses that apply identical or distinct algorithms on data produced by the corresponding simulation. Today, In situ methods are used to analyze large volumes of data generated by scientific simulations at runtime. This work studies the execution of ensemble-based simulations paired with In situ analyses using in-memory staging methods. Because simulations and analyses forming an ensemble typically run concurrently, deploying an ensemble requires efficient co-location-aware strategies, making sure the data flow between simulations and analyses that form an In situ workflow is efficient. Using an ensemble of molecular dynamics In situ workflows with multiple simulations and analyses, we first show that collecting traditional metrics such as makespan, instructions per cycle, memory usage, or cache miss ratio is not sufficient to characterize the complex behaviors of ensembles. Thus, we propose a method to evaluate the performance of ensembles of workflows that captures resource usage (efficiency), resource allocation, and component placement. Experimental results demonstrate that our proposed method can effectively capture the performance of different component placements in an ensemble. By evaluating different co-location scenarios, our performance indicator demonstrates improvements of up to four orders of magnitude when co-locating simulation and coupled analyses within a single computational host.
more »
« less
Human–machine partnerships at the exascale: exploring simulation ensembles through image databases
The explosive growth in supercomputers capacity has changed simulation paradigms. Simulations have shifted from a few lengthy ones to an ensemble of multiple simulations with varying initial conditions or input parameters. Thus, an ensemble consists of large volumes of multi-dimensional data that could go beyond the exascale boundaries. However, the disparity in growth rates between storage capabilities and computing resources results in I/O bottlenecks. This makes it impractical to utilize conventional postprocessing and visualization tools for analyzing such massive simulation ensembles. In situ visualization approaches alleviate I/O constraints by saving predetermined visualizations in image databases during simulation. Nevertheless, the unavailability of output raw data restricts the flexibility of post hoc exploration of in situ approaches. Much research has been conducted to mitigate this limitation, but it falls short when it comes to simultaneously exploring and analyzing parameter and ensemble spaces. In this paper, we propose an expert-in-the-loop visual exploration analytic approach. The proposed approach leverages: feature extraction, deep learning, and human expert–AI collaboration techniques to explore and analyze image-based ensembles. Our approach utilizes local features and deep learning techniques to learn the image features of ensemble members. The extracted features are then combined with simulation input parameters and fed to the visualization pipeline for in-depth exploration and analysis using human expert + AI interaction techniques. We show the effectiveness of our approach using several scientific simulation ensembles.
more »
« less
- PAR ID:
- 10514608
- Publisher / Repository:
- Springer Nature
- Date Published:
- Journal Name:
- Journal of Visualization
- ISSN:
- 1343-8875
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Rectifying AI-generated protein structure ensembles for equilibrium using physics-based computationsRecently, a number of tools have been released that generate ensembles of protein structures based on artificial intelligence (AI) approaches. Although ensembles generated by the tools differ significantly, we demonstrate a computational path to harmonizing the various outputs under a stationary condition using two complementary physics-based approaches. In the first stage, the AI ensemble is used to seed a weighted ensemble (WE) simulation, promoting relaxation toward the steady state. In the second stage, trajectory segments generated by WE are reweighted to steady state using the recently developed RiteWeight (RW) algorithm. We applied this approach to generate an atomically- detailed equilibrium ensemble of unliganded adenylate kinase conformations, starting from ensembles produced by three AI tools: AFSample2, ESMFlow-PDB (trained from PDB structures), and ESMFlow-MD (trained from molecular dynamics simulation data). Dramatic differences in the AI-generated ensembles are largely erased during the WE-RW process, yielding a consistent description of the equilibrium ensemble for a given force field.more » « less
-
Scientific breakthroughs in biomolecular methods and improvements in hardware technology have shifted from a long-running simulation to a large set of shorter simulations running simultaneously, called an ensemble. In an ensemble, simulations are usually coupled with analyses of data produced by the simulations. In situ methods can be used to analyze large volumes of data generated by scientific simulations at runtime (i.e., simulations and analyses are performed concurrently). In this work, we study the execution of ensemble-based simulations paired with in situ analyses using in-memory staging methods. Using an ensemble of molecular dynamics in situ workflows with multiple simulations and analyses, we first show that collecting traditional metrics such as makespan, instructions per cycle, memory usage, or cache miss ratio is not sufficient to characterize complex behaviors of ensembles. We propose a method to evaluate the performance of ensembles of workflows that captures multiple resource usage aspects: resource efficiency, resource allocation, and resource provisioning. Experimental results demonstrate that the proposed method can effectively distinguish the performance of different component placements in an ensemble with up to 32 ensemble members. By evaluating different co-location scenarios, our proposed performance indicators demonstrate benefits of co-locating simulation and coupled analyses within a compute node.more » « less
-
A significant challenge on an exascale computer is the speed at which we compute results exceeds by many orders of magnitude the speed at which we save these results. Therefore the Exascale Computing Project (ECP) ALPINE project focuses on providing exascale-ready visualization solutions including in situ processing. In situ visualization and analysis runs as the simulation is run, on simulations results are they are generated avoiding the need to save entire simulations to storage for later analysis. The ALPINE project made post hoc visualization tools, ParaView and VisIt, exascale ready and developed in situ algorithms and infrastructures. The suite of ALPINE algorithms developed under ECP includes novel approaches to enable automated data analysis and visualization to focus on the most important aspects of the simulation. Many of the algorithms also provide data reduction benefits to meet the I/O challenges at exascale. ALPINE developed a new lightweight in situ infrastructure, Ascent.more » « less
-
Abstract Theoretical models for polycrystalline grain growth are deterministic. However, although computer simulations based on these models reproduce average features and trends, they do not reliably predict experimental growth trajectories for individual grains. Disagreement between experiment and simulation has generally been attributed to shortcomings in the computational instantiation; however, even after several decades of improving the physical bases of computational models, a perfect match has not been achieved. In this study, we examine the sources of uncertainty in simulation and experiment during polycrystalline grain growth. In ensembles of nominally identical molecular dynamics simulations, growth trajectories of individual grains can vary significantly due to discrete events (topological transformations) that have cascading effects on the microstructural ensemble. The type and timing of these microstructural events are extraordinarily sensitive to atomic-scale processes. When we compare ensemble simulation results to experimental outcomes, we find that the ensemble simulation delineates a range of possible outcomes, and the experimental results fall within that range. The implication is that polycrystalline grain growth has a fundamental, aleatoric uncertainty that limits our ability to predict its outcomes. Simulation and experiment can never agree perfectly; however, simulations can predict the range of possible experimental outcomes.more » « less
An official website of the United States government

