Conformal prediction (CP), a distribution-free uncertainty quantification (UQ) framework, reliably provides valid predictive inference for black-box models. CP constructs prediction sets or intervals that contain the true output with a specified probability. However, modern data science’s diverse modalities, along with increasing data and model complexity, challenge traditional CP methods. These developments have spurred novel approaches to address evolving scenarios. This survey reviews the foundational concepts of CP and recent advancements from a data-centric perspective, including applications to structured, unstructured, and dynamic data. We also discuss the challenges and opportunities CP faces in large-scale data and models.
more »
« less
A Modern Theory of Cross-validation Through the Lens of Stability
Modern data analysis and statistical learning are characterized by two defining features: complex data structures and black-box algorithms. The complexity of data structures arises from advanced data collection technologies and data-sharing infrastructures, such as imaging, remote sensing, wearable devices, and genomic sequencing. In parallel, black-box algorithms—particularly those stemming from advances in deep neural networks—have demonstrated remarkable success on modern datasets. This confluence of complex data and opaque models introduces new challenges for uncertainty quantification and statistical inference, a problem we refer to as ``black-box inference''. The difficulty of black-box inference lies in the absence of traditional parametric or nonparametric modeling assumptions, as well as the intractability of the algorithmic behavior underlying many modern estimators. These factors make it difficult to precisely characterize the sampling distribution of estimation errors. A common approach to address this issue is post-hoc randomization, which includes permutation, resampling, sample splitting, cross-validation, and noise injection. When combined with mild assumptions, such as exchangeability in the data-generating process, these methods can yield valid inference and uncertainty quantification. Post-hoc randomization methods have a rich history, ranging from classical techniques like permutation tests, the jackknife, and the bootstrap, to more recent developments such as conformal inference. These approaches typically require minimal knowledge about the underlying data distribution or the inner workings of the estimation procedure. While originally designed for varied purposes, many of these techniques rely, either implicitly or explicitly, on the assumption that the estimation procedure behaves similarly under small perturbations to the data. This idea, now formalized under the concept of \emph{stability}, has become a foundational principle in modern data science. Over the past few decades, stability has emerged as a central research focus in both statistics and machine learning, playing critical roles in areas such as generalization error, data privacy, and adaptive inference. In this article, we investigate one of the most widely used resampling techniques for model comparison and evaluation---cross-validation (CV)---through the lens of stability. We begin by reviewing recent theoretical developments in CV for generalization error estimation and model selection under stability assumptions. We then explore more refined results concerning uncertainty quantification for CV-based risk estimates. By integrating these research directions, we uncover new theoretical insights and methodological tools. Finally, we illustrate their utility across both classical and contemporary topics, including model selection, selective inference, and conformal prediction.
more »
« less
- PAR ID:
- 10688365
- Publisher / Repository:
- Emerald
- Date Published:
- Journal Name:
- Foundations and Trends® in Statistics
- Volume:
- 1
- Issue:
- 3-4
- ISSN:
- 2978-4212
- Page Range / eLocation ID:
- 391 to 548
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Large language models (LLMs) increasingly perform multi-step reasoning, where intermediate claims form implicit directed acyclic graphs whose node correctness is structurally conditioned on their ancestors. This makes factuality uncertainty structural, rather than a trivial accumulation of node-wise errors, and necessitates inferencetime uncertainty quantification over the reasoning structure. While conformal prediction (CP) offers flexible user-specified factuality control, existing work remains post-hoc and cannot intervene during generation. To fill the gap between CP’s flexibility and its post-hoc limitation, we propose an Inference-Time Conformal Reasoning (ITCR) framework that integrates CP directly into reasoning graph generation. ITCR learns a structurelevel factuality uncertainty function that aggregates claim-level factuality signals over reasoning graphs without complex modeling assumptions. We then design the non-conformity score based on graph-level factuality uncertainty and calibrate the conformal threshold to decide when to stop generation. We theoretically show such generation is nested, yielding valid coverage guarantees for factuality control. Experiments over multiple datasets and coverage objectives demonstrate empirically valid coverage. In downstream reasoning tasks, inference-time calibrated graphs yield more accurate generation than post-hoc pruned graphs.more » « less
-
We present a practical guide for the analysis of regression discontinuity (RD) designs in biomedical contexts. We begin by introducing key concepts, assumptions, and estimands within both the continuity‐based framework and the local randomization framework. We then discuss modern estimation and inference methods within both frameworks, including approaches for bandwidth or local neighborhood selection, optimal treatment effect point estimation, and robust bias‐corrected inference methods for uncertainty quantification. We also overview empirical falsification tests that can be used to support key assumptions. Our discussion focuses on two particular features that are relevant in biomedical research: (i) fuzzy RD designs, which often arise when therapeutic treatments are based on clinical guidelines, but patients with scores near the cutoff are treated contrary to the assignment rule; and (ii) RD designs with discrete scores, which are ubiquitous in biomedical applications. We illustrate our discussion with three empirical applications: the effect CD4 guidelines for anti‐retroviral therapy on retention of HIV patients in South Africa, the effect of genetic guidelines for chemotherapy on breast cancer recurrence in the United States, and the effects of age‐based patient cost‐sharing on healthcare utilization in Taiwan. Complete replication materials employing publicly available data and statistical software inPython,RandStataare provided, offering researchers all necessary tools to conduct an RD analysis.more » « less
-
Accurate uncertainty quantification is crucial for making reliable decisions in various supervised learning scenarios, particularly when dealing with complex, multimodal data such as images and text. Current approaches often face notable limitations, including rigid assumptions and limited generalizability, constraining their effectiveness across diverse supervised learning tasks. To overcome these limitations, we introduce Generative Score Inference (GSI), a flexible inference framework capable of constructing statistically valid and informative prediction and confidence sets across a wide range of multimodal learning problems. GSI utilizes synthetic samples generated by deep generative models to approximate conditional score distributions, facilitating precise uncertainty quantification without imposing restrictive assumptions about the data or tasks. We empirically validate GSI’s capabilities through two representative scenarios: hallucination detection in large language models and uncertainty estimation in image captioning. Our method achieves state-of-the-art performance in hallucination detection and robust predictive uncertainty in image captioning, and its performance is positively influenced by the quality of the underlying generative model. These findings underscore the potential of GSI as a versatile inference framework, significantly enhancing uncertainty quantification and trustworthiness in multimodal learning.more » « less
-
Abstract We propose a computationally efficient method to construct nonparametric, heteroscedastic prediction bands for uncertainty quantification, with or without any user-specified predictive model. Our approach provides an alternative to the now-standard conformal prediction for uncertainty quantification, with novel theoretical insights and computational advantages. The data-adaptive prediction band is universally applicable with minimal distributional assumptions, has strong non-asymptotic coverage properties, and is easy to implement using standard convex programs. Our approach can be viewed as a novel variance interpolation with confidence and further leverages techniques from semi-definite programming and sum-of-squares optimization. Theoretical and numerical performances for the proposed approach for uncertainty quantification are analysed.more » « less
An official website of the United States government

