NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

A Fast, Scalable, Universal Approach For Distributed Data Aggregations

https://doi.org/10.1109/BigData50022.2020.9378124

Perera, Niranda; Abeykoon, Vibhatha; Widanage, Chathura; Kamburugamuve, Supun; Kanewala, Thejaka Amila; Wickramasinghe, Pulasthi; Uyar, Ahmet; Maithree, Hasara; Lenadora, Damitha; Fox, Geoffrey (March 2021, 2020 IEEE International Conference on Big Data (Big Data))

Full Text Available
Data Engineering for HPC with Python

https://doi.org/10.1109/PyHPC51966.2020.00007

Abeykoon, Vibhatha; Perera, Niranda; Widanage, Chathura; Kamburugamuve, Supun; Kanewala, Thejaka Amila; Maithree, Hasara; Wickramasinghe, Pulasthi; Uyar, Ahmet; Fox, Geoffrey (November 2020, 2020 IEEE/ACM 9th Workshop on Python for High-Performance and Scientific Computing (PyHPC))
null (Ed.)
Data engineering is becoming an increasingly important part of scientific discoveries with the adoption of deep learning and machine learning. Data engineering deals with a variety of data formats, storage, data extraction, transformation, and data movements. One goal of data engineering is to transform data from original data to vector/matrix/tensor formats accepted by deep learning and machine learning applications. There are many structures such as tables, graphs, and trees to represent data in these data engineering phases. Among them, tables are a versatile and commonly used format to load and process data. In this paper, we present a distributed Python API based on table abstraction for representing and processing data. Unlike existing state-of-the-art data engineering tools written purely in Python, our solution adopts high performance compute kernels in C++, with an in-memory table representation with Cython-based Python bindings. In the core system, we use MPI for distributed memory computations with a data-parallel approach for processing large datasets in HPC clusters.
more » « less
Full Text Available
High Performance Data Engineering Everywhere

https://doi.org/10.1109/SMDS49396.2020.00022

Widanage, Chathura; Perera, Niranda; Abeykoon, Vibhatha; Kamburugamuve, Supun; Kanewala, Thejaka Amila; Maithree, Hasara; Wickramasinghe, Pulasthi; Uyar, Ahmet; Gunduz, Gurhan; Fox, Geoffrey (October 2020, 2020 IEEE International Conference on Smart Data Services (SMDS))
null (Ed.)
The amazing advances being made in the fields of machine and deep learning are a highlight of the Big Data era for both enterprise and research communities. Modern applications require resources beyond a single node's ability to provide. However this is just a small part of the issues facing the overall data processing environment, which must also support a raft of data engineering for pre- and post-data processing, communication, and system integration. An important requirement of data analytics tools is to be able to easily integrate with existing frameworks in a multitude of languages, thereby increasing user productivity and efficiency. All this demands an efficient and highly distributed integrated approach for data processing, yet many of today's popular data analytics tools are unable to satisfy all these requirements at the same time. In this paper we present Cylon, an open-source high performance distributed data processing library that can be seamlessly integrated with existing Big Data and AI/ML frameworks. It is developed with a flexible C++ core on top of a compact data structure and exposes language bindings to C++, Java, and Python. We discuss Cylon's architecture in detail, and reveal how it can be imported as a library to existing applications or operate as a standalone framework. Initial experiments show that Cylon enhances popular tools such as Apache Spark and Dask with major performance improvements for key operations and better component linkages. Finally, we show how its design enables Cylon to be used cross-platform with minimum overhead, which includes popular AI tools such as PyTorch, Tensorflow, and Jupyter notebooks.
more » « less
Full Text Available
Twister2 Cross‐platform resource scheduler for big data

https://doi.org/10.1002/cpe.6502

Uyar, Ahmet; Gunduz, Gurhan; Kamburugamuve, Supun; Wickramasinghe, Pulasthi; Widanage, Chathura; Govindarajan, Kannan; Perera, Niranda; Abeykoon, Vibhatha; Akkas, Selahattin; Fox, Geoffrey (July 2021, Concurrency and Computation: Practice and Experience)

Abstract Twister2 is an open‐source big data hosting environment designed to process both batch and streaming data at scale. Twister2 runs jobs in both high‐performance computing (HPC) and big data clusters. It provides a cross‐platform resource scheduler to run jobs in diverse environments. Twister2 is designed with a layered architecture to support various clusters and big data problems. In this paper, we present the cross‐platform resource scheduler of Twister2. We identify required services and explain implementation details. We present job startup delays for single jobs and multiple concurrent jobs in Kubernetes and OpenMPI clusters. We compare job startup delays for Twister2 and Spark at a Kubernetes cluster. In addition, we compare the performance of terasort algorithm on Kubernetes and bare metal clusters at AWS cloud.
more » « less

Search for: All records