Distributionally Robust Q-Learning

Liu, Zijian; Bai, Qinxun; Blanchet, Jose H.; Dong, Perry; Xu, Wei; Zhou, Zhengqing; Zhou, Zhengyuan

Citation Details

Reinforcement learning (RL) has demonstrated remarkable achievements in simulated environments. However, carrying this success to real environments requires the important attribute of robustness, which the existing RL algorithms often lack as they assume that the future deployment environment is the same as the training environment (i.e. simulator) in which the policy is learned. This assumption often does not hold due to the discrepancy between the simulator and the real environment and, as a result, and hence renders the learned policy fragile when deployed. In this paper, we propose a novel distributionally robust Q-learning algorithm that learns the best policy in the worst distributional perturbation of the environment. Our algorithm first transforms the infinite-dimensional learning problem (since the environment MDP perturbation lies in an infinite-dimensional space) into a finite-dimensional dual problem and subsequently uses a multi-level Monte-Carlo scheme to approximate the dual value using samples from the simulator. Despite the complexity, we show that the resulting distributionally robust Q-learning algorithm asymptotically converges to optimal worst-case policy, thus making it robust to future environment changes. Simulation results further demonstrate its strong empirical robustness. more »

Award ID(s):: 2118199

PAR ID:: 10413185

Author(s) / Creator(s):: Liu, Zijian; Bai, Qinxun; Blanchet, Jose H.; Dong, Perry; Xu, Wei; Zhou, Zhengqing; Zhou, Zhengyuan

Editor(s):: Chaudhuri, Kamalika; Jegelka, Stefanie; Song, Le; Szepesvari, Csaba; Niu, Gang; Sabato, Sivan

Date Published:: 2022-07-01

Journal Name:: Proceedings of Machine Learning Research

ISSN:: 2640-3498

Page Range / eLocation ID:: 13623-13643

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this