Evaluating and Understanding Scheming Propensity in LLM Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Hopman, Mia, Elstner, Jannes, Avramidou, Maria, Prasad, Amritanshu, Lindner, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Consistency Training while Mitigating Obfuscation via Rate Matching
di: Imran, Sohaib, et al.
Pubblicazione: (2026)
di: Imran, Sohaib, et al.
Pubblicazione: (2026)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
di: Wollschläger, Tom, et al.
Pubblicazione: (2025)
di: Wollschläger, Tom, et al.
Pubblicazione: (2025)
Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure
di: Yildirim, Caglar
Pubblicazione: (2026)
di: Yildirim, Caglar
Pubblicazione: (2026)
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
di: Wiedermann-Möller, Jonas, et al.
Pubblicazione: (2026)
di: Wiedermann-Möller, Jonas, et al.
Pubblicazione: (2026)
Combining Cost-Constrained Runtime Monitors for AI Safety
di: Hua, Tim Tian, et al.
Pubblicazione: (2025)
di: Hua, Tim Tian, et al.
Pubblicazione: (2025)
Towards Understanding Specification Gaming in Reasoning Models
di: Nishimura-Gasparian, Kei, et al.
Pubblicazione: (2026)
di: Nishimura-Gasparian, Kei, et al.
Pubblicazione: (2026)
Propensity Inference: Environmental Contributors to LLM Behaviour
di: Järviniemi, Olli, et al.
Pubblicazione: (2026)
di: Järviniemi, Olli, et al.
Pubblicazione: (2026)
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
di: Chen, Jiangjie, et al.
Pubblicazione: (2023)
di: Chen, Jiangjie, et al.
Pubblicazione: (2023)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
di: Zheng, Junhao, et al.
Pubblicazione: (2025)
di: Zheng, Junhao, et al.
Pubblicazione: (2025)
Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
di: Shao, Jiaqi, et al.
Pubblicazione: (2025)
di: Shao, Jiaqi, et al.
Pubblicazione: (2025)
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
di: Luo, Hanjun, et al.
Pubblicazione: (2025)
di: Luo, Hanjun, et al.
Pubblicazione: (2025)
PolicyBank: Evolving Policy Understanding for LLM Agents
di: Choi, Jihye, et al.
Pubblicazione: (2026)
di: Choi, Jihye, et al.
Pubblicazione: (2026)
Evaluating the Propensity of Generative AI for Producing Harmful Disinformation During the 2024 US Election Cycle
di: Schlicht, Erik J
Pubblicazione: (2024)
di: Schlicht, Erik J
Pubblicazione: (2024)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
di: Atinafu, Yonas, et al.
Pubblicazione: (2026)
di: Atinafu, Yonas, et al.
Pubblicazione: (2026)
Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity
di: Yang, Yingxuan, et al.
Pubblicazione: (2026)
di: Yang, Yingxuan, et al.
Pubblicazione: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
The Necessity of a Unified Framework for LLM-Based Agent Evaluation
di: Zhu, Pengyu, et al.
Pubblicazione: (2026)
di: Zhu, Pengyu, et al.
Pubblicazione: (2026)
Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
di: Du, Pengfei
Pubblicazione: (2026)
di: Du, Pengfei
Pubblicazione: (2026)
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
di: Nandi, Subhrangshu, et al.
Pubblicazione: (2025)
di: Nandi, Subhrangshu, et al.
Pubblicazione: (2025)
PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses
di: Badmus, Emmanuel O., et al.
Pubblicazione: (2025)
di: Badmus, Emmanuel O., et al.
Pubblicazione: (2025)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
di: Naik, Akshat, et al.
Pubblicazione: (2025)
di: Naik, Akshat, et al.
Pubblicazione: (2025)
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
di: Jiang, Xinyan, et al.
Pubblicazione: (2026)
di: Jiang, Xinyan, et al.
Pubblicazione: (2026)
RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation
di: Kattamuri, Ashish, et al.
Pubblicazione: (2025)
di: Kattamuri, Ashish, et al.
Pubblicazione: (2025)
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
di: Priyanshu, Aman, et al.
Pubblicazione: (2026)
di: Priyanshu, Aman, et al.
Pubblicazione: (2026)
Discovery of False Data Injection Schemes on Frequency Controllers with Reinforcement Learning
di: Prasad, Romesh, et al.
Pubblicazione: (2024)
di: Prasad, Romesh, et al.
Pubblicazione: (2024)
Evaluation and Benchmarking of LLM Agents: A Survey
di: Mohammadi, Mahmoud, et al.
Pubblicazione: (2025)
di: Mohammadi, Mahmoud, et al.
Pubblicazione: (2025)
MIRAI: Evaluating LLM Agents for Event Forecasting
di: Ye, Chenchen, et al.
Pubblicazione: (2024)
di: Ye, Chenchen, et al.
Pubblicazione: (2024)
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
di: Liu, Ruoqi, et al.
Pubblicazione: (2026)
di: Liu, Ruoqi, et al.
Pubblicazione: (2026)
Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions
di: Xu, Yue, et al.
Pubblicazione: (2026)
di: Xu, Yue, et al.
Pubblicazione: (2026)
GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
di: Li, Yuchen, et al.
Pubblicazione: (2025)
di: Li, Yuchen, et al.
Pubblicazione: (2025)
Evaluating Multi-Turn Bargain Skills in LLM-Based Seller Agent
di: Wang, Issue Yishu, et al.
Pubblicazione: (2025)
di: Wang, Issue Yishu, et al.
Pubblicazione: (2025)
CORE: Full-Path Evaluation of LLM Agents Beyond Final State
di: Michelakis, Panagiotis, et al.
Pubblicazione: (2025)
di: Michelakis, Panagiotis, et al.
Pubblicazione: (2025)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
di: Wu, Yunze, et al.
Pubblicazione: (2025)
di: Wu, Yunze, et al.
Pubblicazione: (2025)
Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
di: Zhao, Cen Mia, et al.
Pubblicazione: (2025)
di: Zhao, Cen Mia, et al.
Pubblicazione: (2025)
Scheming Ability in LLM-to-LLM Strategic Interactions
di: Pham, Thao
Pubblicazione: (2025)
di: Pham, Thao
Pubblicazione: (2025)
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
di: Chaudhary, Manav, et al.
Pubblicazione: (2024)
di: Chaudhary, Manav, et al.
Pubblicazione: (2024)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
di: Choukrani, Omar, et al.
Pubblicazione: (2025)
di: Choukrani, Omar, et al.
Pubblicazione: (2025)
Quantifying the Necessity of Chain of Thought through Opaque Serial Depth
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2026)
di: Brown-Cohen, Jonah, et al.
Pubblicazione: (2026)
The Propensity for Density in Feed-forward Models
di: Schoots, Nandi, et al.
Pubblicazione: (2024)
di: Schoots, Nandi, et al.
Pubblicazione: (2024)
ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping Assistants
di: Wang, Pei, et al.
Pubblicazione: (2026)
di: Wang, Pei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Consistency Training while Mitigating Obfuscation via Rate Matching
di: Imran, Sohaib, et al.
Pubblicazione: (2026) -
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
di: Wollschläger, Tom, et al.
Pubblicazione: (2025) -
Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure
di: Yildirim, Caglar
Pubblicazione: (2026) -
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
di: Wiedermann-Möller, Jonas, et al.
Pubblicazione: (2026) -
Combining Cost-Constrained Runtime Monitors for AI Safety
di: Hua, Tim Tian, et al.
Pubblicazione: (2025)