NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Axolotl: Fairness through Assisted Prompt Rewriting of Large Language Model Outputs

https://doi.org/10.1109/ICKG63256.2024.00017

Ebrahimi, Sana; Chen, Kaiwen; Asudeh, Abolfazl; Das, Gautam; Koudas, Nick (December 2024, IEEE)

Full Text Available
Shapley Values for Explanation in Two-sided Matching Applications

https://doi.org/10.48786/edbt.2024.50

Shetiya, Suraj; Swift, Ian P; Asudeh, Abolfazl; Das, Gautam (January 2024, International Conference on Extending Database Technology (EDBT))
Efficient approximate top-k mutual information based feature selection

https://doi.org/10.1007/s10844-022-00750-4

Salam, Md Abdus; Basu Roy, Senjuti; Das, Gautam (September 2022, Journal of Intelligent Information Systems)

Full Text Available
Fairness-Aware Range Queries for Selecting Unbiased Data

Shetiya, Suraj; Swift, Ian; Asudeh, Abolfazl; Das, Gautam (May 2022, ICDE'22: IEEE international conference on Data Engineering)

We are being constantly judged by automated decision systems that have been widely criticised for being discriminatory and unfair. Since an algorithm is only as good as the data it works with, biases in the data can significantly amplify unfairness issues. In this paper, we take initial steps towards integrating fairness conditions into database query processing and data management systems. Specifically, we focus on selection bias in range queries. We formally define the problem of fairness-aware range queries as obtaining a fair query which is most similar to the user's query. We propose a sub-linear time algorithm for single-predicate range queries and efficient algorithms for multi-predicate range queries. Our empirical evaluation on real and synthetic datasets confirms the effectiveness and efficiency of our proposal.
more » « less
Full Text Available
Fairness-Aware Range Queries for Selecting Unbiased Data

https://doi.org/10.1109/ICDE53745.2022.00111

Shetiya, Suraj; Swift, Ian P.; Asudeh, Abolfazl; Das, Gautam (May 2022, ICDE'22: IEEE international conference on Data Engineering)

Full Text Available
On Finding Rank Regret Representatives

https://doi.org/10.1145/3531054

Asudeh, Abolfazl; Das, Gautam; Jagadish, H. V.; Lu, Shangqi; Nazi, Azade; Tao, Yufei; Zhang, Nan; Zhao, Jianwen (April 2022, ACM Transactions on Database Systems)

Selecting the best items in a dataset is a common task in data exploration. However, the concept of “best” lies in the eyes of the beholder: different users may consider different attributes more important and, hence, arrive at different rankings. Nevertheless, one can remove “dominated” items and create a “representative” subset of the data, comprising the “best items” in it. A Pareto-optimal representative is guaranteed to contain the best item of each possible ranking, but it can be a large portion of data. A much smaller representative can be found if we relax the requirement of including the best item for each user and, instead, just limit the users’ “regret”. Existing work defines regret as the loss in score by limiting consideration to the representative instead of the full dataset, for any chosen ranking function. However, the score is often not a meaningful number, and users may not understand its absolute value. Sometimes small ranges in score can include large fractions of the dataset. In contrast, users do understand the notion of rank ordering. Therefore, we consider items’ positions in the ranked list in defining the regret and propose the rank-regret representative as the minimal subset of the data containing at least one of the top-k of any possible ranking function. This problem is polynomial time solvable in 2D space but is NP-hard on 3 or more dimensions. We design a suite of algorithms to fulfill different purposes, such as whether relaxation is permitted on k, the result size, or both, whether a distribution is known, whether theoretical guarantees or practical efficiency is important, etc. Experiments on real datasets demonstrate that we can efficiently find small subsets with small rank-regrets.
more » « less
Full Text Available
Shahin: Faster Algorithms for Generating Explanations for Multiple Predictions

https://doi.org/10.1145/3448016.3457332

Hasani, Sona; Thirumuruganathan, Saravanan; Koudas, Nick; Das, Gautam (June 2021, ACM SIGMOD)
null (Ed.)
Full Text Available
A Generalized Approach for Reducing Expensive Distance Calls for A Broad Class of Proximity Problems

https://doi.org/10.1145/3448016.3457303

Augustine, Jees; Shetiya, Suraj; Esfandiari, Mohammadreza; Basu Roy, Senjuti; Das, Gautam (June 2021, ACM SIGMOD)
null (Ed.)
Full Text Available
Astrid: accurate selectivity estimation for string predicates using deep learning

https://doi.org/10.14778/3436905.3436907

Shetiya, Suraj; Thirumuruganathan, Saravanan; Koudas, Nick; Das, Gautam (December 2020, Proceedings of the VLDB Endowment)
null (Ed.)
Accurate selectivity estimation for string predicates is a long-standing research challenge in databases. Supporting pattern matching on strings (such as prefix, substring, and suffix) makes this problem much more challenging, thereby necessitating a dedicated study. Traditional approaches often build pruned summary data structures such as tries followed by selectivity estimation using statistical correlations. However, this produces insufficiently accurate cardinality estimates resulting in the selection of sub-optimal plans by the query optimizer. Recently proposed deep learning based approaches leverage techniques from natural language processing such as embeddings to encode the strings and use it to train a model. While this is an improvement over traditional approaches, there is a large scope for improvement. We propose Astrid, a framework for string selectivity estimation that synthesizes ideas from traditional and deep learning based approaches. We make two complementary contributions. First, we propose an embedding algorithm that is query-type (prefix, substring, and suffix) and selectivity aware. Consider three strings 'ab', 'abc' and 'abd' whose prefix frequencies are 1000, 800 and 100 respectively. Our approach would ensure that the embedding for 'ab' is closer to 'abc' than 'abd'. Second, we describe how neural language models could be used for selectivity estimation. While they work well for prefix queries, their performance for substring queries is sub-optimal. We modify the objective function of the neural language model so that it could be used for estimating selectivities of pattern matching queries. We also propose a novel and efficient algorithm for optimizing the new objective function. We conduct extensive experiments over benchmark datasets and show that our proposed approaches achieve state-of-the-art results.
more » « less
Full Text Available
A unified optimization algorithm for solving "regret-minimizing representative" problems

https://doi.org/10.14778/3368289.3368291

Shetiya, Suraj; Asudeh, Abolfazl; Ahmed, Sadia; Das, Gautam (November 2019, Proceedings of the VLDB Endowment)
null (Ed.)
Full Text Available

« Prev Next »

Search for: All records