BATCH: Machine Learning Inference Serving on Serverless Platforms with Adaptive Batching

A. Ali, R. Pinciroli

doi:10.1109/SC41405.2020.00073

Citation Details

BATCH: Machine Learning Inference Serving on Serverless Platforms with Adaptive Batching

Serverless computing is a new pay-per-use cloud service paradigm that automates resource scaling for stateless functions and can potentially facilitate bursty machine learning serving. Batching is critical for latency performance and cost-effectiveness of machine learning inference, but unfortunately it is not supported by existing serverless platforms due to their stateless design. Our experiments show that without batching, machine learning serving cannot reap the benefits of serverless computing. In this paper, we present BATCH, a framework for supporting efficient machine learning serving on serverless platforms. BATCH uses an optimizer to provide inference tail latency guarantees and cost optimization and to enable adaptive batching support. We prototype BATCH atop of AWS Lambda and popular machine learning inference systems. The evaluation verifies the accuracy of the analytic optimizer and demonstrates performance and cost advantages over the state-of-the-art method MArk and the state-of-the-practice tool SageMaker. more »

Award ID(s):: 1838022 1838024 1756013

NSF-PAR ID:: 10206149

Author(s) / Creator(s):: A. Ali, R. Pinciroli

Date Published:: 2020-11-16

Journal Name:: 2020 SC20: International Conference for High Performance Computing, Networking, Storage and Analysis (SC), Atlanta, GA, US, 2020 pp. 972-986. doi: 10.1109/SC41405.2020.00073

Volume:: 1

Page Range / eLocation ID:: 972-986

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.1109/SC41405.2020.00073

More Like this