Guardado en:
| Autores principales: | Hopman, Mia, Elstner, Jannes, Avramidou, Maria, Prasad, Amritanshu, Lindner, David |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.01608 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Consistency Training while Mitigating Obfuscation via Rate Matching
por: Imran, Sohaib, et al.
Publicado: (2026)
por: Imran, Sohaib, et al.
Publicado: (2026)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
por: Wollschläger, Tom, et al.
Publicado: (2025)
por: Wollschläger, Tom, et al.
Publicado: (2025)
Combining Cost-Constrained Runtime Monitors for AI Safety
por: Hua, Tim Tian, et al.
Publicado: (2025)
por: Hua, Tim Tian, et al.
Publicado: (2025)
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
por: Wiedermann-Möller, Jonas, et al.
Publicado: (2026)
por: Wiedermann-Möller, Jonas, et al.
Publicado: (2026)
Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure
por: Yildirim, Caglar
Publicado: (2026)
por: Yildirim, Caglar
Publicado: (2026)
Towards Understanding Specification Gaming in Reasoning Models
por: Nishimura-Gasparian, Kei, et al.
Publicado: (2026)
por: Nishimura-Gasparian, Kei, et al.
Publicado: (2026)
Propensity Inference: Environmental Contributors to LLM Behaviour
por: Järviniemi, Olli, et al.
Publicado: (2026)
por: Järviniemi, Olli, et al.
Publicado: (2026)
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
por: Chen, Jiangjie, et al.
Publicado: (2023)
por: Chen, Jiangjie, et al.
Publicado: (2023)
PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses
por: Badmus, Emmanuel O., et al.
Publicado: (2025)
por: Badmus, Emmanuel O., et al.
Publicado: (2025)
Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
por: Shao, Jiaqi, et al.
Publicado: (2025)
por: Shao, Jiaqi, et al.
Publicado: (2025)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
por: Naik, Akshat, et al.
Publicado: (2025)
por: Naik, Akshat, et al.
Publicado: (2025)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
por: Zheng, Junhao, et al.
Publicado: (2025)
por: Zheng, Junhao, et al.
Publicado: (2025)
PolicyBank: Evolving Policy Understanding for LLM Agents
por: Choi, Jihye, et al.
Publicado: (2026)
por: Choi, Jihye, et al.
Publicado: (2026)
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
por: Luo, Hanjun, et al.
Publicado: (2025)
por: Luo, Hanjun, et al.
Publicado: (2025)
Evaluating the Propensity of Generative AI for Producing Harmful Disinformation During the 2024 US Election Cycle
por: Schlicht, Erik J
Publicado: (2024)
por: Schlicht, Erik J
Publicado: (2024)
Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity
por: Yang, Yingxuan, et al.
Publicado: (2026)
por: Yang, Yingxuan, et al.
Publicado: (2026)
Discovery of False Data Injection Schemes on Frequency Controllers with Reinforcement Learning
por: Prasad, Romesh, et al.
Publicado: (2024)
por: Prasad, Romesh, et al.
Publicado: (2024)
RADAR: Mechanistic Pathways for Detecting Data Contamination in LLM Evaluation
por: Kattamuri, Ashish, et al.
Publicado: (2025)
por: Kattamuri, Ashish, et al.
Publicado: (2025)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
por: Atinafu, Yonas, et al.
Publicado: (2026)
por: Atinafu, Yonas, et al.
Publicado: (2026)
OmniPatch: A Universal Adversarial Patch for ViT-CNN Cross-Architecture Transfer in Semantic Segmentation
por: Aggarwal, Aarush, et al.
Publicado: (2026)
por: Aggarwal, Aarush, et al.
Publicado: (2026)
Quantifying the Necessity of Chain of Thought through Opaque Serial Depth
por: Brown-Cohen, Jonah, et al.
Publicado: (2026)
por: Brown-Cohen, Jonah, et al.
Publicado: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
por: Guo, Zhengkang, et al.
Publicado: (2026)
por: Guo, Zhengkang, et al.
Publicado: (2026)
Scheming Ability in LLM-to-LLM Strategic Interactions
por: Pham, Thao
Publicado: (2025)
por: Pham, Thao
Publicado: (2025)
The Necessity of a Unified Framework for LLM-Based Agent Evaluation
por: Zhu, Pengyu, et al.
Publicado: (2026)
por: Zhu, Pengyu, et al.
Publicado: (2026)
Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
por: Du, Pengfei
Publicado: (2026)
por: Du, Pengfei
Publicado: (2026)
The Propensity for Density in Feed-forward Models
por: Schoots, Nandi, et al.
Publicado: (2024)
por: Schoots, Nandi, et al.
Publicado: (2024)
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
por: Nandi, Subhrangshu, et al.
Publicado: (2025)
por: Nandi, Subhrangshu, et al.
Publicado: (2025)
Evaluation and Benchmarking of LLM Agents: A Survey
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
MIRAI: Evaluating LLM Agents for Event Forecasting
por: Ye, Chenchen, et al.
Publicado: (2024)
por: Ye, Chenchen, et al.
Publicado: (2024)
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
por: Jiang, Xinyan, et al.
Publicado: (2026)
por: Jiang, Xinyan, et al.
Publicado: (2026)
Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
por: Zhao, Cen Mia, et al.
Publicado: (2025)
por: Zhao, Cen Mia, et al.
Publicado: (2025)
Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities
por: Xie, Ziwen, et al.
Publicado: (2026)
por: Xie, Ziwen, et al.
Publicado: (2026)
Beyond Demand Estimation: Consumer Surplus Evaluation via Cumulative Propensity Weights
por: Bian, Zeyu, et al.
Publicado: (2026)
por: Bian, Zeyu, et al.
Publicado: (2026)
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
por: Imran, Mia Mohammad, et al.
Publicado: (2025)
por: Imran, Mia Mohammad, et al.
Publicado: (2025)
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
por: Priyanshu, Aman, et al.
Publicado: (2026)
por: Priyanshu, Aman, et al.
Publicado: (2026)
Gram: Assessing sabotage propensities via automated alignment auditing
por: Lindner, David, et al.
Publicado: (2026)
por: Lindner, David, et al.
Publicado: (2026)
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
por: Chaudhary, Manav, et al.
Publicado: (2024)
por: Chaudhary, Manav, et al.
Publicado: (2024)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
por: Choukrani, Omar, et al.
Publicado: (2025)
por: Choukrani, Omar, et al.
Publicado: (2025)
Survey on Evaluation of LLM-based Agents
por: Yehudai, Asaf, et al.
Publicado: (2025)
por: Yehudai, Asaf, et al.
Publicado: (2025)
ArgMed-Agents: Explainable Clinical Decision Reasoning with LLM Disscusion via Argumentation Schemes
por: Hong, Shengxin, et al.
Publicado: (2024)
por: Hong, Shengxin, et al.
Publicado: (2024)
Ejemplares similares
-
Consistency Training while Mitigating Obfuscation via Rate Matching
por: Imran, Sohaib, et al.
Publicado: (2026) -
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
por: Wollschläger, Tom, et al.
Publicado: (2025) -
Combining Cost-Constrained Runtime Monitors for AI Safety
por: Hua, Tim Tian, et al.
Publicado: (2025) -
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
por: Wiedermann-Möller, Jonas, et al.
Publicado: (2026) -
Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure
por: Yildirim, Caglar
Publicado: (2026)