Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM Interruptions

Levy, Sebastien; Yao, Randolph; Wu, Youjiang; Dang, Yingnong; Huang, Peng; Mu, Zheng; Zhao, Pu; Ramani, Tarun; Govindraju, Naga; Li, Xukun; Lin, Qingwei; Shafriri, Gil Lapid; Chintalapati, Murali

Citation Details

When a failure occurs in production systems, the highest priority is to quickly mitigate it. Despite its importance, failure mitigation is done in a reactive and ad-hoc way: taking some fixed actions only after a severe symptom is observed. For cloud systems, such a strategy is inadequate. In this paper, we propose a preventive and adaptive failure mitigation service, Narya, that is integrated in a production cloud, Microsoft Azure's compute platform. Narya predicts imminent host failures based on multi-layer system signals and then decides smart mitigation actions. The goal is to avert VM failures. Narya's decision engine takes a novel online experimentation approach to continually explore the best mitigation action. Narya further enhances the adaptive decision capability through reinforcement learning. Narya has been running in production for 15 months. It on average reduces VM interruptions by 26% compared to the previous static strategy. more »

Award ID(s):: 1942794

PAR ID:: 10227105

Author(s) / Creator(s):: Levy, Sebastien; Yao, Randolph; Wu, Youjiang; Dang, Yingnong; Huang, Peng; Mu, Zheng; Zhao, Pu; Ramani, Tarun; Govindraju, Naga; Li, Xukun; Lin, Qingwei; Shafriri, Gil Lapid; Chintalapati, Murali

Date Published:: 2020-11-04

Journal Name:: Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this