skip to main content


Search for: All records

Creators/Authors contains: "Nicolae, Bogdan"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Free, publicly-accessible full text available December 20, 2024
  2. Free, publicly-accessible full text available December 18, 2024
  3. Free, publicly-accessible full text available December 20, 2024
  4. Free, publicly-accessible full text available December 18, 2024
  5. Free, publicly-accessible full text available November 12, 2024
  6. Free, publicly-accessible full text available August 7, 2024
  7. With the emergence of versatile storage systems, multi-level checkpointing (MLC) has become a common approach to gain efficiency. However, multi-level checkpoint/restart can cause enormous I/O traffic on HPC systems. To use multilevel checkpointing efficiently, it is important to optimize checkpoint/restart configurations. Current approaches, namely modeling and simulation, are either inaccurate or slow in determining the optimal configuration for a large scale system. In this paper, we show that machine learning models can be used in combination with accurate simulation to determine the optimal checkpoint configurations. We also demonstrate that more advanced techniques such as neural networks can further improve the performance in optimizing checkpoint configurations. 
    more » « less