Optimizing Asynchronous Multi-Level Checkpoint/Restart Configurations with Machine Learning

Dey, Tonmoy; Sato, Kento; Nicolae, Bogdan; Guo, Jian; Domke, Jens; Yu, Weikuan; Cappello, Franck; Mohror, Kathryn

Citation Details

With the emergence of versatile storage systems, multi-level checkpointing (MLC) has become a common approach to gain efficiency. However, multi-level checkpoint/restart can cause enormous I/O traffic on HPC systems. To use multilevel checkpointing efficiently, it is important to optimize checkpoint/restart configurations. Current approaches, namely modeling and simulation, are either inaccurate or slow in determining the optimal configuration for a large scale system. In this paper, we show that machine learning models can be used in combination with accurate simulation to determine the optimal checkpoint configurations. We also demonstrate that more advanced techniques such as neural networks can further improve the performance in optimizing checkpoint configurations. more »

Award ID(s):: 1763547 1744336 1822737 1564647 1561041

PAR ID:: 10156303

Author(s) / Creator(s):: Dey, Tonmoy; Sato, Kento; Nicolae, Bogdan; Guo, Jian; Domke, Jens; Yu, Weikuan; Cappello, Franck; Mohror, Kathryn

Date Published:: 2020-05-01

Journal Name:: The IEEE International Workshop on High-Performance Storage

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this