Active Advantage-Aligned Online Reinforcement Learning with Offline Data

Liu, X; Le, HT; Chen, S; Stevens, R; Yang, Z; Walter, MR; Chen, Y

Citation Details

This content will become publicly available on July 13, 2026

Active Advantage-Aligned Online Reinforcement Learning with Offline Data

Online reinforcement learning (RL) enhances policies through direct interactions with the environment, but faces challenges related to sample efficiency. In contrast, offline RL leverages extensive pre-collected data to learn policies, but often produces suboptimal results due to limited data coverage. Recent efforts integrate offline and online RL in order to harness the advantages of both approaches. However, effectively combining online and offline RL remains challenging due to issues that include catastrophic forgetting, lack of robustness to data quality and limited sample efficiency in data utilization. In an effort to address these challenges, we introduce A3RL, which incorporates a novel confidence aware Active Advantage Aligned (A3) sampling strategy that dynamically prioritizes data aligned with the policy's evolving needs from both online and offline sources, optimizing policy improvement. Moreover, we provide theoretical insights into the effectiveness of our active sampling strategy and conduct diverse empirical experiments and ablation studies, demonstrating that our method outperforms competing online RL techniques that leverage offline data. Our code will be publicly available at:this https URL. more »

Award ID(s):: 2332475

PAR ID:: 10621891

Author(s) / Creator(s):: Liu, X; Le, HT; Chen, S; Stevens, R; Yang, Z; Walter, MR; Chen, Y

Publisher / Repository:: Exploration in AI Today Workshop at ICML (ExAI), July 2025.

Date Published:: 2025-07-13

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
This content will become publicly available on July 13, 2026
Workshop Report:
The DOI is not currently available.

More Like this