Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Free, publicly-accessible full text available November 4, 2026
-
Abstract In a variety of applications, including nonparametric instrumental variable (NPIV) analysis, proximal causal inference under unmeasured confounding, and analysis of missing-not-at-random data with shadow variables, we are interested in inference on a continuous linear functional (e.g. average causal effects) of nuisance functions (e.g. NPIV regression) defined by conditional moment restrictions. These nuisance functions are often weakly identified, meaning the moment restrictions are ill-posed and may admit multiple solutions. This paper proposes a novel condition for the functional to be strongly identified (amenable to n) rate asymptotically normal estimation) even when the nuisance function remains weakly identified. The condition implies the existence of debiasing nuisance functions. We propose penalized minimax estimators for both the primary and debiasing nuisance functions. These estimators accommodate flexible function classes and, crucially, converge to fixed limits determined by the penalization, irrespective of the nuisances' identifiability. We use these penalized estimators to construct a debiased functional estimator and prove its asymptotic normality under general high-level conditions, leading to valid confidence intervals. Our method is illustrated in partially linear proximal causal inference and instrumental variable regression problems.more » « lessFree, publicly-accessible full text available December 3, 2026
-
In this paper, we prove that Distributional Reinforcement Learning (DistRL), which learns the return distribution, can obtain second-order bounds in both online and offline RL in general settings with function approximation. Second-order bounds are instance-dependent bounds that scale with the variance of return, which we prove are tighter than the previously known small-loss bounds of distributional RL. To the best of our knowledge, our results are the first second-order bounds for low-rank MDPs and for offline RL. When specializing to contextual bandits (one-step RL problem), we show that a distributional learning based optimism algorithm achieves a second-order worst-case regret bound, and a second-order gap dependent bound, simultaneously. We also empirically demonstrate the benefit of DistRL in contextual bandits on real-world datasets. We highlight that our analysis with DistRL is relatively simple, follows the general framework of optimism in the face of uncertainty and does not require weighted regression. Our results suggest that DistRL is a promising framework for obtaining second-order bounds in general RL settings, thus further reinforcing the benefits of DistRL.more » « less
-
In this paper, we prove that Distributional Re- inforcement Learning (DistRL), which learns the return distribution, can obtain second-order bounds in both online and offline RL in general settings with function approximation. Second- order bounds are instance-dependent bounds that scale with the variance of return, which we prove are tighter than the previously known small-loss bounds of distributional RL. To the best of our knowledge, our results are the first second-order bounds for low-rank MDPs and for offline RL. When specializing to contextual bandits (one-step RL problem), we show that a distributional learn- ing based optimism algorithm achieves a second- order worst-case regret bound, and a second-order gap dependent bound, simultaneously. We also empirically demonstrate the benefit of DistRL in contextual bandits on real-world datasets. We highlight that our analysis with DistRL is rela- tively simple, follows the general framework of optimism in the face of uncertainty and does not require weighted regression. Our results suggest that DistRL is a promising framework for obtain- ing second-order bounds in general RL settings, thus further reinforcing the benefits of DistRL.more » « less
An official website of the United States government

Full Text Available