Mainlining databases: supporting fast transactional workloads on universal columnar data file formats

Li, Tianyu; Butrovich, Matthew; Ngom, Amadou; Lim, Wan Shen; McKinney, Wes; Pavlo, Andrew

doi:10.14778/3436905.3436913

Citation Details

Mainlining databases: supporting fast transactional workloads on universal columnar data file formats

The proliferation of modern data processing tools has given rise to open-source columnar data formats. These formats help organizations avoid repeated conversion of data to a new format for each application. However, these formats are read-only, and organizations must use a heavy-weight transformation process to load data from on-line transactional processing (OLTP) systems. As a result, DBMSs often fail to take advantage of full network bandwidth when transferring data. We aim to reduce or even eliminate this overhead by developing a storage architecture for in-memory database management systems (DBMSs) that is aware of the eventual usage of its data and emits columnar storage blocks in a universal open-source format. We introduce relaxations to common analytical data formats to efficiently update records and rely on a lightweight transformation process to convert blocks to a read-optimized layout when they are cold. We also describe how to access data from third-party analytical tools with minimal serialization overhead. We implemented our storage engine based on the Apache Arrow format and integrated it into the NoisePage DBMS to evaluate our work. Our experiments show that our approach achieves comparable performance with dedicated OLTP DBMSs while enabling orders-of-magnitude faster data exports to external data science and machine learning tools than existing methods. more »

Award ID(s):: 1822933 1846158

PAR ID:: 10312178

Author(s) / Creator(s):: Li, Tianyu; Butrovich, Matthew; Ngom, Amadou; Lim, Wan Shen; McKinney, Wes; Pavlo, Andrew

Date Published:: 2020-12-01

Journal Name:: Proceedings of the VLDB Endowment

Volume:: 14

Issue:: 4

ISSN:: 2150-8097

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
https://doi.org/10.14778/3436905.3436913

More Like this