Improving scalability and reliability of MPI-agnostic transparent checkpointing for production workloads at NERSC

Chouhan, Prashant Singh; Khetawat, Harsh; Resnik, Neil; Jain, Twinkle; Garg, Rohan; Cooperman, Gene; Hartman-Baker, Rebecca; Zhao, Zhengji

Citation Details

Checkpoint/restart (C/R) provides fault-tolerant computing capability, enables long running applications, and provides scheduling flexibility for computing centers to support diverse workloads with different priority. It is therefore vital to get transparent C/R capability working at NERSC. MANA, by Garg et. al., is a transparent checkpointing tool that has been selected due to its MPI-agnostic and network-agnostic approach. However, originally written as a proof-of-concept code, MANA was not ready to use with NERSC's diverse production workloads, which are dominated by MPI and hybrid MPI+OpenMP applications. In this talk, we present ongoing work at NERSC to enable MANA for NERSC's production workloads, including fixing bugs that were exposed by the top applications at NERSC, adding new features to address system changes, evaluating C/R overhead at scale, etc. The lessons learned from making MANA production-ready for HPC applications will be useful for C/R tool developers, supercomputing centers and HPC end-users alike. more »

Award ID(s):: 1740218

PAR ID:: 10319924

Author(s) / Creator(s):: Chouhan, Prashant Singh; Khetawat, Harsh; Resnik, Neil; Jain, Twinkle; Garg, Rohan; Cooperman, Gene; Hartman-Baker, Rebecca; Zhao, Zhengji

Date Published:: 2021-02-01

Journal Name:: First International Symposium on Checkpointing for Supercomputing (SuperCheck21)

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this