NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Stochastic SketchRefine: Scaling In-Database Decision-Making under Uncertainty to Millions of Tuples

Haque, Riddho R; Mai, Anh L; Brucato, Matteo; Abouzied, Azza; Haas, Peter J; Meliou, Alexandra (September 2025, Proceedings of the VLDB Endowment)

Decision making under uncertainty often requires choosing packages, or bags of tuples, that collectively optimize expected outcomes while limiting risks. Processing Stochastic Package Queries (SPQs) involves solving very large optimization problems on uncertain data. Monte Carlo methods create numerous scenarios, or sample realizations of the stochastic attributes of all the tuples, and generate packages with optimal objective values across these scenarios. The number of scenarios needed for accurate approximation---and hence the size of the optimization problem when using prior methods---increases with variance in the data, and the search space of the optimization problem increases exponentially with the number of tuples in the relation. Existing solvers take hours to process SPQs on large relations containing stochastic attributes with high variance. Besides enriching the SPaQL language to capture a broader class of risk specifications, we make two fundamental contributions toward scalable SPQ processing. First, we propose risk-constraint linearization (RCL), which converts SPQs into Integer Linear Programs (ILPs) whose size is independent of the number of scenarios used. Solving these ILPs gives us feasible and near-optimal packages. Second, we propose Stochastic Sketch Refine, a divide and conquer framework that breaks down a large stochastic optimization problem into subproblems involving smaller subsets of tuples. Our experiments show that, together, RCL and Stochastic Sketch Refine produce high-quality packages in orders of magnitude lower runtime than the state of the art.
more » « less
Free, publicly-accessible full text available September 1, 2026
Scaling Package Queries to a Billion Tuples via Hierarchical Partitioning and Customized Optimization

https://doi.org/10.14778/3641204.3641222

Mai, Anh L; Wang, Pengyu; Abouzied, Azza; Brucato, Matteo; Haas, Peter J; Meliou, Alexandra (January 2024, Proceedings of the VLDB Endowment)

A package query returns a package---a multiset of tuples---that maximizes or minimizes a linear objective function subject to linear constraints, thereby enabling in-database decision support. Prior work has established the equivalence of package queries to Integer Linear Programs (ILPs) and developed the SketchRefine algorithm for package query processing. While this algorithm was an important first step toward supporting prescriptive analytics scalably inside a relational database, it struggles when the data size grows beyond a few hundred million tuples or when the constraints become very tight. In this paper, we present Progressive Shading, a novel algorithm for processing package queries that can scale efficiently to billions of tuples and gracefully handle tight constraints. Progressive Shading solves a sequence of optimization problems over a hierarchy of relations, each resulting from an ever-finer partitioning of the original tuples into homogeneous groups until the original relation is obtained. This strategy avoids the premature discarding of high-quality tuples that can occur with SketchRefine. Our novel partitioning scheme, Dynamic Low Variance, can handle very large relations with multiple attributes and can dynamically adapt to both concentrated and spread-out sets of attribute values, provably outperforming traditional partitioning schemes such as kd-tree. We further optimize our system by replacing our off-the-shelf optimization software with customized ILP and LP solvers, called Dual Reducer and Parallel Dual Simplex respectively, that are highly accurate and orders of magnitude faster.
more » « less
Full Text Available
Through the Data Management Lens: Experimental Analysis and Evaluation of Fair Classification

https://doi.org/10.1145/3514221.3517841

Islam, Maliha Tashfia; Fariha, Anna; Meliou, Alexandra; Salimi, Babak (June 2022, Proceedings of the 2022 International Conference on Management of Data (SIGMOD))

Full Text Available
DataPrism: Exposing Disconnect between Data and Systems

https://doi.org/10.1145/3514221.3517864

Galhotra, Sainyam; Fariha, Anna; Lourenço, Raoni; Freire, Juliana; Meliou, Alexandra; Srivastava, Divesh (June 2022, Proceedings of the 2022 International Conference on Management of Data (SIGMOD))

Full Text Available
Improved Approximation and Scalability for Fair Max-Min Diversification

https://doi.org/10.4230/LIPIcs.ICDT.2022.7

Addanki, Raghavendra; McGregor, Andrew; Meliou, Alexandra; Moumoulidou, Zafeiria (March 2022, 25th International Conference on Database Theory (ICDT))

Full Text Available
In-Database Decision Support: Opportunities and Challenges

Azza Abouzied; Peter J. Haas; Alexandra Meliou (January 2022, A Quarterly bulletin of the IEEE Computer Society Technical Committee on Database Engineering)

Full Text Available
Trends in Explanations: Understanding and Debugging Data-driven Systems

https://doi.org/10.1561/1900000074

Glavic, Boris; Meliou, Alexandra; Roy, Sudeepa (January 2021, Foundations and Trends® in Databases)
null (Ed.)
Full Text Available
Diverse Data Selection under Fairness Constraints

https://doi.org/10.4230/LIPIcs.ICDT.2021.13

Moumoulidou, Zafeiria; McGregor, Andrew (January 2021, International Conference on Database Theory)

Diversity is an important principle in data selection and summarization, facility location, and recommendation systems. Our work focuses on maximizing diversity in data selection, while offering fairness guarantees. In particular, we offer the first study that augments the Max-Min diversification objective with fairness constraints. More specifically, given a universe 𝒰 of n elements that can be partitioned into m disjoint groups, we aim to retrieve a k-sized subset that maximizes the pairwise minimum distance within the set (diversity) and contains a pre-specified k_i number of elements from each group i (fairness). We show that this problem is NP-complete even in metric spaces, and we propose three novel algorithms, linear in n, that provide strong theoretical approximation guarantees for different values of m and k. Finally, we extend our algorithms and analysis to the case where groups can be overlapping.
more » « less
Full Text Available
SUBSUME: A Dataset for Subjective Summary Extraction from Wikipedia Documents

https://doi.org/10.18653/v1/2021.newsum-1.14

Yadav, Nishant; Brucato, Matteo; Fariha, Anna; Youngquist, Oscar; Killingback, Julian; Meliou, Alexandra; Haas, Peter (January 2021, New Frontiers in Summarization workshop (at EMNLP 2021))

Full Text Available
sPaQLTooLs: a stochastic package query interface for scalable constrained optimization

https://doi.org/10.14778/3415478.3415499

Brucato, Matteo; Mannino, Miro; Abouzied, Azza; Haas, Peter J.; Meliou, Alexandra (August 2020, Proceedings of the VLDB Endowment)

Full Text Available

« Prev Next »

Search for: All records