Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance Guarantees

Panaganti, Kishan; Kalathil, Dileep

Citation Details

This paper addresses the problem of model-free reinforcement learning for Robust Markov Decision Process (RMDP) with large state spaces. The goal of the RMDP framework is to find a policy that is robust against the parameter uncertainties due to the mismatch between the simulator model and real-world settings. We first propose the Ro- bust Least Squares Policy Evaluation algorithm, which is a multi-step online model-free learning algorithm for policy evaluation. We prove the convergence of this algorithm using stochastic approximation techniques. We then propose Robust Least Squares Policy Iteration (RLSPI) algorithm for learning the optimal robust policy. We also give a general weighted Euclidean norm bound on the error (closeness to optimality) of the resulting policy. Finally, we demonstrate the performance of our RLSPI algorithm on some standard bench- mark problems. more »

Award ID(s):: 2045783

PAR ID:: 10327511

Author(s) / Creator(s):: Panaganti, Kishan; Kalathil, Dileep

Date Published:: 2021-07-01

Journal Name:: International Conference on Machine Learning (ICML)

Page Range / eLocation ID:: 511-520

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this