Visual Transfer for Reinforcement Learning via Wasserstein Domain Confusion

Roy, J; Konidaris, G.D.

Citation Details

We introduce Wasserstein Adversarial Proximal Policy Optimization (WAPPO), a novel algorithm for visual transfer in Reinforcement Learning that explicitly learns to align the distributions of extracted features between a source and target task. WAPPO approximates and minimizes the Wasserstein-1 distance between the distributions of features from source and target domains via a novel Wasserstein Confusion objective. WAPPO outperforms the prior state-of-the-art in visual transfer and successfully transfers policies across Visual Cartpole and both the easy and hard settings of of 16 OpenAI Procgen environments. more »

Award ID(s):: 1717569

PAR ID:: 10310141

Author(s) / Creator(s):: Roy, J; Konidaris, G.D.

Date Published:: 2021-02-01

Journal Name:: Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
The DOI is not currently available.

More Like this