Gespeichert in:
| Hauptverfasser: | Li, Christy, CH-Wang, Sky, Peng, Andi, Bobu, Andreea |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2604.18847 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
von: Damani, Mehul, et al.
Veröffentlicht: (2024)
von: Damani, Mehul, et al.
Veröffentlicht: (2024)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
Aligning Robot and Human Representations
von: Bobu, Andreea, et al.
Veröffentlicht: (2023)
von: Bobu, Andreea, et al.
Veröffentlicht: (2023)
Do Androids Know They're Only Dreaming of Electric Sheep?
von: CH-Wang, Sky, et al.
Veröffentlicht: (2023)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2023)
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
von: Deshpande, Darshan, et al.
Veröffentlicht: (2024)
von: Deshpande, Darshan, et al.
Veröffentlicht: (2024)
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
von: Jones, Jaylen, et al.
Veröffentlicht: (2026)
von: Jones, Jaylen, et al.
Veröffentlicht: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
Robots That Know What to Ask: Recovering Misaligned Rewards through Targeted Explanations
von: Merker, Helena, et al.
Veröffentlicht: (2026)
von: Merker, Helena, et al.
Veröffentlicht: (2026)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
von: Zhu, Shenzhe
Veröffentlicht: (2025)
von: Zhu, Shenzhe
Veröffentlicht: (2025)
InforME: Improving Informativeness of Abstractive Text Summarization With Informative Attention Guided by Named Entity Salience
von: Shen, Jianbin, et al.
Veröffentlicht: (2025)
von: Shen, Jianbin, et al.
Veröffentlicht: (2025)
Scaling Agents for Computer Use
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2025)
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2025)
Preference-Conditioned Language-Guided Abstraction
von: Peng, Andi, et al.
Veröffentlicht: (2024)
von: Peng, Andi, et al.
Veröffentlicht: (2024)
Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language
von: Hwang, Minyoung, et al.
Veröffentlicht: (2025)
von: Hwang, Minyoung, et al.
Veröffentlicht: (2025)
Efficient Agent Training for Computer Use
von: He, Yanheng, et al.
Veröffentlicht: (2025)
von: He, Yanheng, et al.
Veröffentlicht: (2025)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2026)
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2026)
QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents
von: Nader, Jordan Abi, et al.
Veröffentlicht: (2025)
von: Nader, Jordan Abi, et al.
Veröffentlicht: (2025)
Training Computer Use Agents to Assess the Usability of Graphical User Interfaces
von: Gao, Alice, et al.
Veröffentlicht: (2026)
von: Gao, Alice, et al.
Veröffentlicht: (2026)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
von: Miles, Roy, et al.
Veröffentlicht: (2026)
von: Miles, Roy, et al.
Veröffentlicht: (2026)
Stress-Testing Model Specs Reveals Character Differences among Language Models
von: Zhang, Jifan, et al.
Veröffentlicht: (2025)
von: Zhang, Jifan, et al.
Veröffentlicht: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
von: Kim, Heehwan, et al.
Veröffentlicht: (2025)
von: Kim, Heehwan, et al.
Veröffentlicht: (2025)
VideoAgentTrek: Computer Use Pretraining from Unlabeled Videos
von: Lu, Dunjie, et al.
Veröffentlicht: (2025)
von: Lu, Dunjie, et al.
Veröffentlicht: (2025)
Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
von: Agashe, Saaket, et al.
Veröffentlicht: (2025)
von: Agashe, Saaket, et al.
Veröffentlicht: (2025)
Value of Information: A Framework for Human-Agent Communication
von: Dong, Yijiang River, et al.
Veröffentlicht: (2026)
von: Dong, Yijiang River, et al.
Veröffentlicht: (2026)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
von: Mohamadi, Alireza, et al.
Veröffentlicht: (2025)
von: Mohamadi, Alireza, et al.
Veröffentlicht: (2025)
MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
von: Li, Binxu, et al.
Veröffentlicht: (2024)
von: Li, Binxu, et al.
Veröffentlicht: (2024)
Agent S: An Open Agentic Framework that Uses Computers Like a Human
von: Agashe, Saaket, et al.
Veröffentlicht: (2024)
von: Agashe, Saaket, et al.
Veröffentlicht: (2024)
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
von: Li, Zhigen, et al.
Veröffentlicht: (2024)
von: Li, Zhigen, et al.
Veröffentlicht: (2024)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
Towards Comprehensive Detection of Chinese Harmful Memes
von: Lu, Junyu, et al.
Veröffentlicht: (2024)
von: Lu, Junyu, et al.
Veröffentlicht: (2024)
StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation
von: Zheng, Huawei, et al.
Veröffentlicht: (2026)
von: Zheng, Huawei, et al.
Veröffentlicht: (2026)
Structsum Generation for Faster Text Comprehension
von: Jain, Parag, et al.
Veröffentlicht: (2024)
von: Jain, Parag, et al.
Veröffentlicht: (2024)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
von: Guo, Xuehang, et al.
Veröffentlicht: (2025)
von: Guo, Xuehang, et al.
Veröffentlicht: (2025)
Milestone-Guided Policy Learning for Long-Horizon Language Agents
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
von: Damani, Mehul, et al.
Veröffentlicht: (2024) -
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2025) -
Aligning Robot and Human Representations
von: Bobu, Andreea, et al.
Veröffentlicht: (2023) -
Do Androids Know They're Only Dreaming of Electric Sheep?
von: CH-Wang, Sky, et al.
Veröffentlicht: (2023) -
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
von: Deshpande, Darshan, et al.
Veröffentlicht: (2024)