When will my ML Job finish? Toward providing Completion Time Estimates through Predictability-Centric Scheduling

Faisal, Abdullah; Martin, Noah; Bashir, Hafiz; Lamelas, Swaminathan; Dogar, Fahad

Citation Details

In this paper, we make a case for providing job completion time estimates to GPU cluster users, similar to providing the delivery date of a package or arrival time of a booked ride. Our analysis reveals that providing predictability can come at the expense of performance and fairness. Existing GPU schedulers optimize for extreme points in the trade-off space, making them either extremely unpredictable or impractical. To address this challenge, we present PCS, a new scheduling framework that aims to provide predictability while balancing other traditional objectives. The key idea behind PCS is to use Weighted-Fair-Queueing (WFQ) and find a suitable configuration of different WFQ parameters (e.g., queue weights) that meets specific goals for predictability. It uses a simulation-aided search strategy to efficiently discover WFQ configurations that lie around the Pareto front of the trade-off space between these objectives. We implement and evaluate PCS in the context of scheduling ML training workloads on GPUs. Our evaluation, on a small-scale GPU testbed and larger-scale simulations, shows that PCS can provide accurate completion time estimates while marginally compromising on performance and fairness. more »

Award ID(s):: 2106797

PAR ID:: 10548868

Author(s) / Creator(s):: Faisal, Abdullah; Martin, Noah; Bashir, Hafiz; Lamelas, Swaminathan; Dogar, Fahad

Publisher / Repository:: Usenix OSDI

Date Published:: 2024-07-10

ISBN:: 978-1-939133-40-3

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this