Human-Guided Harm Recovery for Computer Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Christy, CH-Wang, Sky, Peng, Andi, Bobu, Andreea |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024)
by: Damani, Mehul, et al.
Published: (2024)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2025)
by: Ma, Rachel, et al.
Published: (2025)
Do Androids Know They're Only Dreaming of Electric Sheep?
by: CH-Wang, Sky, et al.
Published: (2023)
by: CH-Wang, Sky, et al.
Published: (2023)
Aligning Robot and Human Representations
by: Bobu, Andreea, et al.
Published: (2023)
by: Bobu, Andreea, et al.
Published: (2023)
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
by: Deshpande, Darshan, et al.
Published: (2024)
by: Deshpande, Darshan, et al.
Published: (2024)
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
by: CH-Wang, Sky, et al.
Published: (2025)
by: CH-Wang, Sky, et al.
Published: (2025)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
by: Jones, Jaylen, et al.
Published: (2026)
by: Jones, Jaylen, et al.
Published: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
by: Zhu, Shenzhe
Published: (2025)
by: Zhu, Shenzhe
Published: (2025)
Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
by: Merker, Helena, et al.
Published: (2026)
by: Merker, Helena, et al.
Published: (2026)
InforME: Improving Informativeness of Abstractive Text Summarization With Informative Attention Guided by Named Entity Salience
by: Shen, Jianbin, et al.
Published: (2025)
by: Shen, Jianbin, et al.
Published: (2025)
Scaling Agents for Computer Use
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
by: Sharshar, Ahmed, et al.
Published: (2026)
by: Sharshar, Ahmed, et al.
Published: (2026)
Training Computer Use Agents to Assess the Usability of Graphical User Interfaces
by: Gao, Alice, et al.
Published: (2026)
by: Gao, Alice, et al.
Published: (2026)
Efficient Agent Training for Computer Use
by: He, Yanheng, et al.
Published: (2025)
by: He, Yanheng, et al.
Published: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
by: Kim, Heehwan, et al.
Published: (2025)
by: Kim, Heehwan, et al.
Published: (2025)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
by: Feng, Yuming, et al.
Published: (2026)
by: Feng, Yuming, et al.
Published: (2026)
Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
by: Miles, Roy, et al.
Published: (2026)
by: Miles, Roy, et al.
Published: (2026)
VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
by: Lu, Dunjie, et al.
Published: (2025)
by: Lu, Dunjie, et al.
Published: (2025)
MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
by: Li, Binxu, et al.
Published: (2024)
by: Li, Binxu, et al.
Published: (2024)
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
by: Li, Zhigen, et al.
Published: (2024)
by: Li, Zhigen, et al.
Published: (2024)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
by: Mohamadi, Alireza, et al.
Published: (2025)
by: Mohamadi, Alireza, et al.
Published: (2025)
Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language
by: Hwang, Minyoung, et al.
Published: (2025)
by: Hwang, Minyoung, et al.
Published: (2025)
Stress-Testing Model Specs Reveals Character Differences among Language Models
by: Zhang, Jifan, et al.
Published: (2025)
by: Zhang, Jifan, et al.
Published: (2025)
StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation
by: Zheng, Huawei, et al.
Published: (2026)
by: Zheng, Huawei, et al.
Published: (2026)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Towards Comprehensive Detection of Chinese Harmful Memes
by: Lu, Junyu, et al.
Published: (2024)
by: Lu, Junyu, et al.
Published: (2024)
Agent S: An Open Agentic Framework that Uses Computers Like a Human
by: Agashe, Saaket, et al.
Published: (2024)
by: Agashe, Saaket, et al.
Published: (2024)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
by: Schoene, Annika M, et al.
Published: (2025)
by: Schoene, Annika M, et al.
Published: (2025)
Preference-Conditioned Language-Guided Abstraction
by: Peng, Andi, et al.
Published: (2024)
by: Peng, Andi, et al.
Published: (2024)
Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
by: Agashe, Saaket, et al.
Published: (2025)
by: Agashe, Saaket, et al.
Published: (2025)
Milestone-Guided Policy Learning for Long-Horizon Language Agents
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents
by: Nader, Jordan Abi, et al.
Published: (2025)
by: Nader, Jordan Abi, et al.
Published: (2025)
FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration
by: Jung, Dongwon, et al.
Published: (2025)
by: Jung, Dongwon, et al.
Published: (2025)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
by: Mekky, Ali, et al.
Published: (2025)
by: Mekky, Ali, et al.
Published: (2025)
TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
by: Chen, Yongchao, et al.
Published: (2025)
by: Chen, Yongchao, et al.
Published: (2025)
Structsum Generation for Faster Text Comprehension
by: Jain, Parag, et al.
Published: (2024)
by: Jain, Parag, et al.
Published: (2024)
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
by: Guo, Xuehang, et al.
Published: (2025)
by: Guo, Xuehang, et al.
Published: (2025)
Similar Items
-
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024) -
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2025) -
Do Androids Know They're Only Dreaming of Electric Sheep?
by: CH-Wang, Sky, et al.
Published: (2023) -
Aligning Robot and Human Representations
by: Bobu, Andreea, et al.
Published: (2023) -
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
by: Deshpande, Darshan, et al.
Published: (2024)