Deep RC: A Scalable Data Engineering and Deep Learning Pipeline

Sarker, Arup; Alsaadi, Aymen; Halpern, Alexander; Tangella, Prabhath; Titov, Mikhail; Perera, Niranda; Staylor, Mills; Laszewski, Gregor von; Jha, Shantenu; Fox, Geoffrey

Citation Details

This content will become publicly available on June 3, 2026

Deep RC: A Scalable Data Engineering and Deep Learning Pipeline

Significant obstacles exist in scientific domains including genetics, climate modeling, and astronomy due to the management, preprocess, and training on complicated data for deep learning. Even while several large-scale solutions offer distributed execution environments, open-source alternatives that integrate scalable runtime tools, deep learning and data frameworks on high-performance computing platforms remain crucial for accessibility and flexibility. In this paper, we introduce Deep Radical-Cylon(RC), a heterogeneous runtime system that combines data engineering, deep learning frameworks, and workflow engines across several HPC environments, including cloud and supercomputing infrastructures. Deep RC supports heterogeneous systems with accelerators, allows the usage of communication libraries like \texttt{MPI}, \texttt{GLOO} and \texttt{NCCL} across multi-node setups, and facilitates parallel and distributed deep learning pipelines by utilizing Radical Pilot as a task execution framework. By attaining an end-to-end pipeline including preprocessing, model training, and postprocessing with 11 neural forecasting models (PyTorch) and hydrology models (TensorFlow) under identical resource conditions, the system reduces 3.28 and 75.9 seconds, respectively. The design of Deep RC guarantees the smooth integration of scalable data frameworks, such as Cylon, with deep learning processes, exhibiting strong performance on cloud platforms and scientific HPC systems. By offering a flexible, high-performance solution for resource-intensive applications, this method closes the gap between data preprocessing, model training, and postprocessing. more »

Award ID(s):: 2212550 2504401

PAR ID:: 10651658

Author(s) / Creator(s):: Sarker, Arup; Alsaadi, Aymen; Halpern, Alexander; Tangella, Prabhath; Titov, Mikhail; Perera, Niranda; Staylor, Mills; Laszewski, Gregor von; Jha, Shantenu; Fox, Geoffrey

Publisher / Repository:: Springer. JSSPP 2025: Job Scheduling Strategies for Parallel Processing

Date Published:: 2025-06-03

Format(s):: Medium: X

Location:: Milan, Italy

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
This content will become publicly available on June 3, 2026
Conference Paper:
The DOI is not currently available.

More Like this