FUDJ: Flexible User-Defined Distributed Joins

Sevim, Akil; Eldawy, Ahmed; Carman, E Preston; Carey, Michael J; Tsotras, Vassilis J

doi:10.1109/ICDE60146.2024.00320

Citation Details

FUDJ: Flexible User-Defined Distributed Joins

Join operations are crucial in data analysis, but can suffer inefficiency with large datasets and complex non-equality-based conditions. Optimized join algorithms have gained traction in database research to address these challenges. One popular choice for implementing join algorithms is distributed data processing frameworks, e.g., Hadoop and Spark, but each implementation is highly tailored for specific query types. As a result, they do not address join queries that involve diverse and complex conditions since they are not integrated into a holistic query optimization engine like in DBMSs. On the other hand, implementing new join algorithms on a DBMS from scratch requires substantial effort and expertise. This paper introduces FUDJ, Flexible User-defined Distributed Joins, a framework for complex distributed join algorithms. The key idea of FUDJ is to allow developers to realize new distributed join algorithms into the database without delving into the database internals. As shown, an algorithm implemented in FUDJ is up to an order of magnitude faster than existing user-defined implementations with an order of magnitude fewer lines of code. more »

Award ID(s):: 1954644 1954962 1924694

PAR ID:: 10546045

Author(s) / Creator(s):: Sevim, Akil; Eldawy, Ahmed; Carman, E Preston; Carey, Michael J; Tsotras, Vassilis J

Publisher / Repository:: IEEE

Date Published:: 2024-05-13

ISBN:: 979-8-3503-1715-2

Page Range / eLocation ID:: 4194 to 4207

Format(s):: Medium: X

Location:: Utrecht, Netherlands

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.1109/ICDE60146.2024.00320

More Like this