Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries. While current benchmarks evaluate general-purpose performance~(e.g., WebArena) or safety against malicious actions~(e.g., SafeArena), no existing framework assesses an agent's ability to successfully execute user-facing website security and privacy tasks, such as managing cookie preferences, configuring privacy-sensitive account settings, or revoking inactive sessions.To address this gap, we introduce WebSP-Eval, an evaluation framework for measuring web agent performance on website security and privacy tasks. WebSP-Eval comprises 1) a manually crafted task dataset of 200 task instances across 28 websites; 2) a robust agentic system supporting account and initial state management across runs using a custom Google Chrome extension; and 3) an automated evaluator. We evaluate a total of 8 web agent instantiations using state-of-the-art multimodal large language models, conducting a fine-grained analysis across websites, task categories, and UI elements. Our evaluation reveals that current models suffer from limited autonomous exploration capabilities to reliably solve website security and privacy tasks, and struggle with specific task categories and websites. Crucially, we identify stateful UI elements are a primary reason for agent failure, with toggles causing more than 45% task failure across many models.more » « lessFree, publicly-accessible full text available October 1, 2027
-
Free, publicly-accessible full text available June 20, 2027
-
Open-Source Software (OSS) is susceptible to supply chain attacks and vulnerabilities introduced by malicious actors. Recent examples include a backdoor enabling remote-code execution discovered in the XZ Utils project and the deliberate addition of Use-After-Free bugs into the Linux kernel. These vulnerabilities often consist of subtle, well-crafted changes designed to evade a small group of maintainers, and can even be injected by malicious maintainers who have spent years building trust. The gem5 simulator faces similar risks, with potentially far-reaching consequences: vulnerabilities introduced into gem5 may propagate into future hardware designs built on its models, especially given its widespread use across academia, industry, and national labs for early-stage design exploration. Although gem5 incorporates a robust testing infrastructure, the codebase remains difficult to maintain due to its breadth of domains and steady influx of hundreds of commits each year. Subtle vulnerabilities like Spectre or Meltdown may be introduced inadvertently or with malicious intent. Scaling the project places pressure on the review process managed by around 20 volunteer maintainers, who are not trained to detect security threats. Thus, there is a need for tools that vet commits and flag potential vulnerabilities before integration into the main branch. We develop a Commit-Code Inconsistency Detection System using Large Language Models (LLMs), which have demonstrated strong capabilities in code reasoning. Our system identifies inconsistent commit message-code pairs and acts as a first step in detecting vulnerable code in OSSs.more » « lessFree, publicly-accessible full text available June 27, 2027
-
Free, publicly-accessible full text available June 25, 2027
-
Free, publicly-accessible full text available November 19, 2026
An official website of the United States government

Full Text Available