RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Zeyi, Jones, Jaylen, Jiang, Linxi, Ning, Yuting, Fosler-Lussier, Eric, Su, Yu, Lin, Zhiqiang, Sun, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
by: Jones, Jaylen, et al.
Published: (2024)
by: Jones, Jaylen, et al.
Published: (2024)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
by: Jones, Jaylen, et al.
Published: (2026)
by: Jones, Jaylen, et al.
Published: (2026)
AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts
by: Kumar, Vishal, et al.
Published: (2024)
by: Kumar, Vishal, et al.
Published: (2024)
Music on the Move
by: Fosler-Lussier, Danielle
Published: (2020)
by: Fosler-Lussier, Danielle
Published: (2020)
Improving Transducer-Based Spoken Language Understanding with Self-Conditioned CTC and Knowledge Transfer
by: Sunder, Vishal, et al.
Published: (2025)
by: Sunder, Vishal, et al.
Published: (2025)
Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers
by: Serai, Prashant, et al.
Published: (2024)
by: Serai, Prashant, et al.
Published: (2024)
End-to-End Diarization utilizing Attractor Deep Clustering
by: Palzer, David, et al.
Published: (2025)
by: Palzer, David, et al.
Published: (2025)
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
by: Palzer, David, et al.
Published: (2025)
by: Palzer, David, et al.
Published: (2025)
When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents
by: Ning, Yuting, et al.
Published: (2026)
by: Ning, Yuting, et al.
Published: (2026)
AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
by: Liao, Zeyi, et al.
Published: (2024)
by: Liao, Zeyi, et al.
Published: (2024)
VISTA: Verification In Sequential Turn-based Assessment
by: Lewis, Ashley, et al.
Published: (2025)
by: Lewis, Ashley, et al.
Published: (2025)
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
by: Chun, Jiyun, et al.
Published: (2026)
by: Chun, Jiyun, et al.
Published: (2026)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
by: Ginjala, Srishti, et al.
Published: (2026)
by: Ginjala, Srishti, et al.
Published: (2026)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
by: Wang, Bowen, et al.
Published: (2026)
by: Wang, Bowen, et al.
Published: (2026)
Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents
by: Xue, Tianci, et al.
Published: (2026)
by: Xue, Tianci, et al.
Published: (2026)
UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
by: Yang, Yuhao, et al.
Published: (2025)
by: Yang, Yuhao, et al.
Published: (2025)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
by: Xu, Chejian, et al.
Published: (2024)
by: Xu, Chejian, et al.
Published: (2024)
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
by: Mo, Lingbo, et al.
Published: (2024)
by: Mo, Lingbo, et al.
Published: (2024)
OpenCUA: Open Foundations for Computer-Use Agents
by: Wang, Xinyuan, et al.
Published: (2025)
by: Wang, Xinyuan, et al.
Published: (2025)
PRO-CUA: Process-Reward Optimization for Computer Use Agents
by: He, Yifei, et al.
Published: (2026)
by: He, Yifei, et al.
Published: (2026)
LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
by: Mei, Kai, et al.
Published: (2025)
by: Mei, Kai, et al.
Published: (2025)
Web Verbs: Typed Abstractions for Reliable Task Composition on the Agentic Web
by: Jiang, Linxi, et al.
Published: (2026)
by: Jiang, Linxi, et al.
Published: (2026)
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
by: Jiang, Linxi, et al.
Published: (2026)
by: Jiang, Linxi, et al.
Published: (2026)
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
by: Liu, Zhaoyang, et al.
Published: (2025)
by: Liu, Zhaoyang, et al.
Published: (2025)
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
by: Jian, Xiangru, et al.
Published: (2026)
by: Jian, Xiangru, et al.
Published: (2026)
Autonomous Adversary: Red-Teaming in the age of LLM
by: Mamun, Mohammad, et al.
Published: (2026)
by: Mamun, Mohammad, et al.
Published: (2026)
WebGuard: Building a Generalizable Guardrail for Web Agents
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
On the Adversarial Risk of Test Time Adaptation: An Investigation into Realistic Test-Time Data Poisoning
by: Su, Yongyi, et al.
Published: (2024)
by: Su, Yongyi, et al.
Published: (2024)
WebArena: A Realistic Web Environment for Building Autonomous Agents
by: Zhou, Shuyan, et al.
Published: (2023)
by: Zhou, Shuyan, et al.
Published: (2023)
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
by: Hu, Xuhao, et al.
Published: (2026)
by: Hu, Xuhao, et al.
Published: (2026)
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
A Non-autoregressive Model for Joint STT and TTS
by: Sunder, Vishal, et al.
Published: (2025)
by: Sunder, Vishal, et al.
Published: (2025)
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
by: Bai, Hao, et al.
Published: (2026)
by: Bai, Hao, et al.
Published: (2026)
EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
by: Xue, Taofeng, et al.
Published: (2026)
by: Xue, Taofeng, et al.
Published: (2026)
Red Teaming AI Red Teaming
by: Majumdar, Subhabrata, et al.
Published: (2025)
by: Majumdar, Subhabrata, et al.
Published: (2025)
Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation
by: Morasso, Cristian, et al.
Published: (2026)
by: Morasso, Cristian, et al.
Published: (2026)
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments
by: Liu, Zefang, et al.
Published: (2025)
by: Liu, Zefang, et al.
Published: (2025)
AttributionBench: How Hard is Automatic Attribution Evaluation?
by: Li, Yifei, et al.
Published: (2024)
by: Li, Yifei, et al.
Published: (2024)
CUA-Skill: Develop Skills for Computer Using Agent
by: Chen, Tianyi, et al.
Published: (2026)
by: Chen, Tianyi, et al.
Published: (2026)
Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Similar Items
-
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
by: Jones, Jaylen, et al.
Published: (2024) -
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
by: Jones, Jaylen, et al.
Published: (2026) -
AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts
by: Kumar, Vishal, et al.
Published: (2024) -
Music on the Move
by: Fosler-Lussier, Danielle
Published: (2020) -
Improving Transducer-Based Spoken Language Understanding with Self-Conditioned CTC and Knowledge Transfer
by: Sunder, Vishal, et al.
Published: (2025)