Signal Temporal Logic-Guided Apprenticeship Learning

Puranic, Aniruddh G; Deshmukh, Jyotirmoy V; Nikolaidis, Stefanos

doi:10.1109/IROS58592.2024.10801924

Citation Details

Signal Temporal Logic-Guided Apprenticeship Learning

Apprenticeship learning crucially depends on effectively learning rewards, and hence control policies from user demonstrations. Of particular difficulty is the setting where the desired task consists of a number of sub-goals with temporal dependencies. The quality of inferred rewards and hence policies are typically limited by the quality of demonstrations, and poor inference of these can lead to undesirable outcomes. In this paper, we show how temporal logic specifications that describe high level task objectives, are encoded in a graph to define a temporal-based metric that reasons about behaviors of demonstrators and the learner agent to improve the quality of inferred rewards and policies. Through experiments on a diverse set of robot manipulator simulations, we show how our framework overcomes the drawbacks of prior literature by drastically improving the number of demonstrations required to learn a control policy. more »

Award ID(s):: 1932620 1837131

PAR ID:: 10564797

Author(s) / Creator(s):: Puranic, Aniruddh G; Deshmukh, Jyotirmoy V; Nikolaidis, Stefanos

Publisher / Repository:: IEEE

Date Published:: 2024-10-14

ISBN:: 979-8-3503-7770-5

Page Range / eLocation ID:: 11147 to 11154

Format(s):: Medium: X

Location:: Abu Dhabi, United Arab Emirates

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Conference Paper:
https://doi.org/10.1109/IROS58592.2024.10801924

More Like this