RT-ZooKeeper: Taming the Recovery Latency of a Coordination Service

Li, Haoran; Lu, Chenyang; Gill, Christopher D.

doi:10.1145/3477034

Citation Details

RT-ZooKeeper: Taming the Recovery Latency of a Coordination Service

Fault-tolerant coordination services have been widely used in distributed applications in cloud environments. Recent years have witnessed the emergence of time-sensitive applications deployed in edge computing environments, which introduces both challenges and opportunities for coordination services. On one hand, coordination services must recover from failures in a timely manner. On the other hand, edge computing employs local networked platforms that can be exploited to achieve timely recovery. In this work, we first identify the limitations of the leader election and recovery protocols underlying Apache ZooKeeper, the prevailing open-source coordination service. To reduce recovery latency from leader failures, we then design RT-Zookeeper with a set of novel features including a fast-convergence election protocol, a quorum channel notification mechanism, and a distributed epoch persistence protocol. We have implemented RT-Zookeeper based on ZooKeeper version 3.5.8. Empirical evaluation shows that RT-ZooKeeper achieves 91% reduction in maximum recovery latency in comparison to ZooKeeper. Furthermore, a case study demonstrates that fast failure recovery in RT-ZooKeeper can benefit a common messaging service like Kafka in terms of message latency. more »

Award ID(s):: 1646579

NSF-PAR ID:: 10391327

Author(s) / Creator(s):: Li, Haoran; Lu, Chenyang; Gill, Christopher D.

Date Published:: 2021-10-31

Journal Name:: ACM Transactions on Embedded Computing Systems

Volume:: 20

Issue:: 5s

ISSN:: 1539-9087

Page Range / eLocation ID:: 1 to 22

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
https://doi.org/10.1145/3477034

More Like this