HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Xiaoxue, Jiang, Penghao, Li, Kaixin, Huang, Zhiyong, Du, Xiaoning, Jiang, Jiaojiao, Xing, Zhenchang, Sun, Jiamou, Zhuo, Terry Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Identifying and Mitigating API Misuse in Large Language Models
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
From Code to Courtroom: LLMs as the New Software Judges
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
PTMPicker: Facilitating Efficient Pretrained Model Selection for Application Developers
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Do Chase Your Tail! Missing Key Aspects Augmentation in Textual Vulnerability Descriptions of Long-tail Software through Feature Inference
by: Han, Linyi, et al.
Published: (2024)
by: Han, Linyi, et al.
Published: (2024)
Domain-constrained Synthesis of Inconsistent Key Aspects in Textual Vulnerability Descriptions
by: Han, Linyi, et al.
Published: (2025)
by: Han, Linyi, et al.
Published: (2025)
An Empirical Study of Vulnerabilities in Python Packages and Their Detection
by: Quan, Haowei, et al.
Published: (2025)
by: Quan, Haowei, et al.
Published: (2025)
To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack
by: Zhuo, Terry Yue, et al.
Published: (2026)
by: Zhuo, Terry Yue, et al.
Published: (2026)
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
by: Jiang, Linxi, et al.
Published: (2026)
by: Jiang, Linxi, et al.
Published: (2026)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
by: Thaman, Kunvar
Published: (2026)
by: Thaman, Kunvar
Published: (2026)
Editorial for the special issue on software refactoring: Application breadth and technical depth
by: Zhenchang Xing
Published: (2024)
by: Zhenchang Xing
Published: (2024)
Favia: Forensic Agent for Vulnerability-fix Identification and Analysis
by: Storhaug, André, et al.
Published: (2026)
by: Storhaug, André, et al.
Published: (2026)
Unveiling the Tricks: Automated Detection of Dark Patterns in Mobile Applications
by: Chen, Jieshan, et al.
Published: (2023)
by: Chen, Jieshan, et al.
Published: (2023)
Refactoring to Pythonic Idioms: A Hybrid Knowledge-Driven Approach Leveraging Large Language Models
by: Zhang, Zejun, et al.
Published: (2024)
by: Zhang, Zejun, et al.
Published: (2024)
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
LLMAID: Identifying AI Capabilities in Android Apps with LLMs
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
by: Zou, Wei, et al.
Published: (2026)
by: Zou, Wei, et al.
Published: (2026)
LayerPlexRank: Exploring Node Centrality and Layer Influence through Algebraic Connectivity in Multiplex Networks
by: Ren, Hao, et al.
Published: (2024)
by: Ren, Hao, et al.
Published: (2024)
Explore-Construct-Filter: An Automated Framework for Rich and Reliable API Knowledge Graph Construction
by: Sun, Yanbang, et al.
Published: (2025)
by: Sun, Yanbang, et al.
Published: (2025)
Bridging Solidity Evolution Gaps: An LLM-Enhanced Approach for Smart Contract Compilation Error Resolution
by: Ye, Likai, et al.
Published: (2025)
by: Ye, Likai, et al.
Published: (2025)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
by: Yang, Ruozhao, et al.
Published: (2026)
by: Yang, Ruozhao, et al.
Published: (2026)
WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
by: Fang, Tianqing, et al.
Published: (2025)
by: Fang, Tianqing, et al.
Published: (2025)
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
by: Kong, Fanheng, et al.
Published: (2026)
by: Kong, Fanheng, et al.
Published: (2026)
Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model
by: Wu, Tianyi, et al.
Published: (2026)
by: Wu, Tianyi, et al.
Published: (2026)
ICE-Score: Instructing Large Language Models to Evaluate Code
by: Zhuo, Terry Yue
Published: (2023)
by: Zhuo, Terry Yue
Published: (2023)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
by: Ding, Junchen, et al.
Published: (2025)
by: Ding, Junchen, et al.
Published: (2025)
Are Latent Vulnerabilities Hidden Gems for Software Vulnerability Prediction? An Empirical Study
by: Le, Triet H. M., et al.
Published: (2024)
by: Le, Triet H. M., et al.
Published: (2024)
A Systematization of Security Vulnerabilities in Computer Use Agents
by: Jones, Daniel, et al.
Published: (2025)
by: Jones, Daniel, et al.
Published: (2025)
Dipole-Obstructed Cooper Pairing: Theory and Application to $j=3/2$ Superconductors
by: Zhu, Penghao, et al.
Published: (2024)
by: Zhu, Penghao, et al.
Published: (2024)
R-WoM: Retrieval-augmented World Model For Computer-use Agents
by: Mei, Kai, et al.
Published: (2025)
by: Mei, Kai, et al.
Published: (2025)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
Influence Robustness of Nodes in Multiplex Networks against Attacks
by: Ma, Boqian, et al.
Published: (2023)
by: Ma, Boqian, et al.
Published: (2023)
From Exploration to Revelation: Detecting Dark Patterns in Mobile Apps
by: Chen, Jieshan, et al.
Published: (2024)
by: Chen, Jieshan, et al.
Published: (2024)
Rethinking Broken Object Level Authorization Attacks Under Zero Trust Principle
by: Wu, Anbin, et al.
Published: (2025)
by: Wu, Anbin, et al.
Published: (2025)
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
by: He, Hongliang, et al.
Published: (2024)
by: He, Hongliang, et al.
Published: (2024)
Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents
by: Ma, SHengjie, et al.
Published: (2025)
by: Ma, SHengjie, et al.
Published: (2025)
A Topology-aware Graph Coarsening Framework for Continual Graph Learning
by: Han, Xiaoxue, et al.
Published: (2024)
by: Han, Xiaoxue, et al.
Published: (2024)
A^3-CodGen: A Repository-Level Code Generation Framework for Code Reuse with Local-Aware, Global-Aware, and Third-Party-Library-Aware
by: Liao, Dianshu, et al.
Published: (2023)
by: Liao, Dianshu, et al.
Published: (2023)
Similar Items
-
Identifying and Mitigating API Misuse in Large Language Models
by: Zhuo, Terry Yue, et al.
Published: (2025) -
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025) -
From Code to Courtroom: LLMs as the New Software Judges
by: He, Junda, et al.
Published: (2025) -
PTMPicker: Facilitating Efficient Pretrained Model Selection for Application Developers
by: Liu, Pei, et al.
Published: (2025) -
Do Chase Your Tail! Missing Key Aspects Augmentation in Textual Vulnerability Descriptions of Long-tail Software through Feature Inference
by: Han, Linyi, et al.
Published: (2024)