Navigating the Unexpected Realities of Big Data Transfers in a Cloud-based World

Rivera, Sergio; Griffioen, James; Fei, Zongming; Hayashida, Mami; Shi, Pinyi; Chitre, Bhushan; Chappell, Jacob; Song, Yongwook; Pike, Lowell; Carpenter, Charles; Nasir, Hussamuddin

doi:10.1145/3219104.3229276

Citation Details

Navigating the Unexpected Realities of Big Data Transfers in a Cloud-based World

The emergence of big data has created new challenges for researchers transmitting big data sets across campus networks to local (HPC) cloud resources, or over wide area networks to public cloud services. Unlike conventional HPC systems where the network is carefully architected (e.g., a high speed local interconnect, or a wide area connection between Data Transfer Nodes), today's big data communication often occurs over shared network infrastructures with many external and uncontrolled factors influencing performance. This paper describes our efforts to understand and characterize the performance of various big data transfer tools such as rclone, cyberduck, and other provider-specific CLI tools when moving data to/from public and private cloud resources. We analyze the various parameter settings available on each of these tools and their impact on performance. Our experimental results give insights into the performance of cloud providers and transfer tools, and provide guidance for parameter settings when using cloud transfer tools. We also explore performance when coming from HPC DTN nodes as well as researcher machines located deep in the campus network, and show that emerging SDN approaches such as the VIP Lanes system can deliver excellent performance even from researchers' machines. more »

Award ID(s):: 1642134 1541380 1541426

PAR ID:: 10073416

Author(s) / Creator(s):: Rivera, Sergio; Griffioen, James; Fei, Zongming; Hayashida, Mami; Shi, Pinyi; Chitre, Bhushan; Chappell, Jacob; Song, Yongwook; Pike, Lowell; Carpenter, Charles; Nasir, Hussamuddin

Date Published:: 2018-07-22

Journal Name:: PEARC '18 Proceedings of the Practice and Experience on Advanced Research Computing

Volume:: 22

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.1145/3219104.3229276

More Like this