Guardagent: Safeguard LLM agents via knowledge-enabled reasoning

Xiang, Zhen; Zheng, Linzhi; Li, Yanjie; Hong, Junyuan; Li, Qinbin; Xie, Han; Zhang, Jiawei; Xiong, Zidi; Xie, Chulin; Bastian, Nathaniel D

Citation Details

The rapid advancement of large language model (LLM) agents has raised new concerns regarding their safety and security, which cannot be addressed by traditional textual-harm-focused LLM guardrails. We propose GuardAgent, the first guardrail agent to protect other agents by checking whether the agent actions satisfy safety guard requests. Specifically, GuardAgent first analyzes the safety guard requests to generate a task plan, and then converts this plan into guardrail code for execution. In both steps, an LLM is utilized as the reasoning component, supplemented by in-context demonstrations retrieved from a memory module storing information from previous tasks. GuardAgent can understand different safety guard requests and provide reliable code-based guardrails with high flexibility and low operational overhead. In addition, we propose two novel benchmarks: EICU-AC benchmark to assess the access control for healthcare agents and Mind2Web-SC benchmark to evaluate the safety regulations for web agents. We show that GuardAgent effectively moderates the violation actions for two types of agents on these two benchmarks with over 98% and 83% guardrail accuracies, respectively. more »

Award ID(s):: 2229876

PAR ID:: 10661454

Author(s) / Creator(s):: Xiang, Zhen; Zheng, Linzhi; Li, Yanjie; Hong, Junyuan; Li, Qinbin; Xie, Han; Zhang, Jiawei; Xiong, Zidi; Xie, Chulin; Bastian, Nathaniel D

Publisher / Repository:: ICML 2025 Workshop on Computer Use Agents

Date Published:: 2025-07-13

Volume:: 267

Format(s):: Medium: X

Location:: Vancouver, Canada

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript
Conference Paper:
The DOI is not currently available.

More Like this