Analyzing Probabilistic Methods for Evaluating Agent Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Højmark, Axel, Pimpale, Govind, Panickssery, Arjun, Hobbhahn, Marius, Scheurer, Jérémy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
Large Language Models Often Know When They Are Being Evaluated
by: Needham, Joe, et al.
Published: (2025)
by: Needham, Joe, et al.
Published: (2025)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024)
by: Meinke, Alexander, et al.
Published: (2024)
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
by: Lermen, Simon, et al.
Published: (2024)
by: Lermen, Simon, et al.
Published: (2024)
Training Deliberative Monitors for Black-Box Scheming Detection
by: Sinha, Aditya, et al.
Published: (2026)
by: Sinha, Aditya, et al.
Published: (2026)
Technical Report: Evaluating Goal Drift in Language Model Agents
by: Arike, Rauno, et al.
Published: (2025)
by: Arike, Rauno, et al.
Published: (2025)
LLM Evaluators Recognize and Favor Their Own Generations
by: Panickssery, Arjun, et al.
Published: (2024)
by: Panickssery, Arjun, et al.
Published: (2024)
A Benchmark for Scalable Oversight Protocols
by: Sudhir, Abhimanyu Pallavi, et al.
Published: (2025)
by: Sudhir, Abhimanyu Pallavi, et al.
Published: (2025)
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025)
by: Schoen, Bronson, et al.
Published: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024)
by: Laine, Rudolf, et al.
Published: (2024)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
by: Ackerman, Christopher, et al.
Published: (2024)
by: Ackerman, Christopher, et al.
Published: (2024)
TracrBench: Generating Interpretability Testbeds with Large Language Models
by: Thurnherr, Hannes, et al.
Published: (2024)
by: Thurnherr, Hannes, et al.
Published: (2024)
Mitigating Many-Shot Jailbreaking
by: Ackerman, Christopher M., et al.
Published: (2025)
by: Ackerman, Christopher M., et al.
Published: (2025)
Memory-Augmented Agent Training for Business Document Understanding
by: Liu, Jiale, et al.
Published: (2024)
by: Liu, Jiale, et al.
Published: (2024)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
by: Ball, Sarah, et al.
Published: (2024)
by: Ball, Sarah, et al.
Published: (2024)
Unlocking Cognitive Capabilities and Analyzing the Perception-Logic Trade-off
by: Zhang, Longyin, et al.
Published: (2026)
by: Zhang, Longyin, et al.
Published: (2026)
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Think Smart, Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning
by: Chatterji, Satchit, et al.
Published: (2024)
by: Chatterji, Satchit, et al.
Published: (2024)
Discovering and Learning Probabilistic Models of Black-Box AI Capabilities
by: Bramblett, Daniel, et al.
Published: (2025)
by: Bramblett, Daniel, et al.
Published: (2025)
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
by: Li, Xiangyi, et al.
Published: (2026)
by: Li, Xiangyi, et al.
Published: (2026)
REBUS: A Robust Evaluation Benchmark of Understanding Symbols
by: Gritsevskiy, Andrew, et al.
Published: (2024)
by: Gritsevskiy, Andrew, et al.
Published: (2024)
Will we run out of data? Limits of LLM scaling based on human-generated data
by: Villalobos, Pablo, et al.
Published: (2022)
by: Villalobos, Pablo, et al.
Published: (2022)
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
by: Black, Sid, et al.
Published: (2025)
by: Black, Sid, et al.
Published: (2025)
Log analysis is necessary for credible evaluation of AI agents
by: Kirgis, Peter, et al.
Published: (2026)
by: Kirgis, Peter, et al.
Published: (2026)
Intentional Deception as Controllable Capability in LLM Agents
by: Starace, Jason, et al.
Published: (2026)
by: Starace, Jason, et al.
Published: (2026)
Evaluating Environments Using Exploratory Agents
by: Khaleque, Bobby, et al.
Published: (2024)
by: Khaleque, Bobby, et al.
Published: (2024)
Evaluating Developmental Cognition Capabilities of LLMs
by: Xiao, Xiao, et al.
Published: (2026)
by: Xiao, Xiao, et al.
Published: (2026)
Tracking Capabilities for Safer Agents
by: Odersky, Martin, et al.
Published: (2026)
by: Odersky, Martin, et al.
Published: (2026)
Analyzing and Internalizing Complex Policy Documents for LLM Agents
by: Liu, Jiateng, et al.
Published: (2025)
by: Liu, Jiateng, et al.
Published: (2025)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
by: Olson, Matthew Lyle, et al.
Published: (2026)
by: Olson, Matthew Lyle, et al.
Published: (2026)
Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation
by: Sudhakar, Arjun Vaithilingam
Published: (2025)
by: Sudhakar, Arjun Vaithilingam
Published: (2025)
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
by: Backlund, Axel, et al.
Published: (2025)
by: Backlund, Axel, et al.
Published: (2025)
A Conceptual Framework for AI Capability Evaluations
by: Carro, María Victoria, et al.
Published: (2025)
by: Carro, María Victoria, et al.
Published: (2025)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
by: Turk, Matt
Published: (2026)
by: Turk, Matt
Published: (2026)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
Operand Quant: A Single-Agent Architecture for Autonomous Machine Learning Engineering
by: Sahney, Arjun, et al.
Published: (2025)
by: Sahney, Arjun, et al.
Published: (2025)
CADMAS-CTX: Contextual Capability Calibration for Multi-Agent Delegation
by: Qiao, Chuhan
Published: (2026)
by: Qiao, Chuhan
Published: (2026)
Responsibility-aware Strategic Reasoning in Probabilistic Multi-Agent Systems
by: Mu, Chunyan, et al.
Published: (2024)
by: Mu, Chunyan, et al.
Published: (2024)
Similar Items
-
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025) -
Large Language Models Often Know When They Are Being Evaluated
by: Needham, Joe, et al.
Published: (2025) -
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023) -
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024) -
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
by: Lermen, Simon, et al.
Published: (2024)