Data-Driven Insight Synthesis for Multi-Dimensional Data

Xing, Junjie; Wang, Xinyu; Jagadish, H V

doi:10.14778/3641204.3641211

Citation Details

Data-Driven Insight Synthesis for Multi-Dimensional Data

Exploratory data analysis can uncover interesting data insights from data. Current methods utilize interestingness measures designed based on system designers' perspectives, thus inherently restricting the insights to their defined scope. These systems, consequently, may not adequately represent a broader range of user interests. Furthermore, most existing approaches that formulate interestingness measure are rule-based, which makes them inevitably brittle and often requires holistic re-design when new user needs are discovered. This paper presents a data-driven technique for deriving an interestingness measure that learns from annotated data. We further develop an innovative annotation algorithm that significantly reduces the annotation cost, and an insight synthesis algorithm based on the Markov Chain Monte Carlo method for efficient discovery of interesting insights. We consolidate these ideas into a system. Our experimental outcomes and user studies demonstrate that DAISY can effectively discover a broad range of interesting insights, thereby substantially advancing the current state-of-the-art. more »

Award ID(s):: 2312931 2106176

PAR ID:: 10535252

Author(s) / Creator(s):: Xing, Junjie; Wang, Xinyu; Jagadish, H V

Publisher / Repository:: VLDB

Date Published:: 2024-01-01

Journal Name:: Proceedings of the VLDB Endowment

Volume:: 17

Issue:: 5

ISSN:: 2150-8097

Page Range / eLocation ID:: 1007 to 1019

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
https://doi.org/10.14778/3641204.3641211

More Like this