Analyzing Probabilistic Methods for Evaluating Agent Capabilities
Fuente:
arXiv
Guardado en:
| Autores principales: | Højmark, Axel, Pimpale, Govind, Panickssery, Arjun, Hobbhahn, Marius, Scheurer, Jérémy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Forecasting Frontier Language Model Agent Capabilities
por: Pimpale, Govind, et al.
Publicado: (2025)
por: Pimpale, Govind, et al.
Publicado: (2025)
Large Language Models Often Know When They Are Being Evaluated
por: Needham, Joe, et al.
Publicado: (2025)
por: Needham, Joe, et al.
Publicado: (2025)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
por: Scheurer, Jérémy, et al.
Publicado: (2023)
por: Scheurer, Jérémy, et al.
Publicado: (2023)
Frontier Models are Capable of In-context Scheming
por: Meinke, Alexander, et al.
Publicado: (2024)
por: Meinke, Alexander, et al.
Publicado: (2024)
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
por: Lermen, Simon, et al.
Publicado: (2024)
por: Lermen, Simon, et al.
Publicado: (2024)
Training Deliberative Monitors for Black-Box Scheming Detection
por: Sinha, Aditya, et al.
Publicado: (2026)
por: Sinha, Aditya, et al.
Publicado: (2026)
Technical Report: Evaluating Goal Drift in Language Model Agents
por: Arike, Rauno, et al.
Publicado: (2025)
por: Arike, Rauno, et al.
Publicado: (2025)
LLM Evaluators Recognize and Favor Their Own Generations
por: Panickssery, Arjun, et al.
Publicado: (2024)
por: Panickssery, Arjun, et al.
Publicado: (2024)
A Benchmark for Scalable Oversight Protocols
por: Sudhir, Abhimanyu Pallavi, et al.
Publicado: (2025)
por: Sudhir, Abhimanyu Pallavi, et al.
Publicado: (2025)
Stress Testing Deliberative Alignment for Anti-Scheming Training
por: Schoen, Bronson, et al.
Publicado: (2025)
por: Schoen, Bronson, et al.
Publicado: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
por: Laine, Rudolf, et al.
Publicado: (2024)
por: Laine, Rudolf, et al.
Publicado: (2024)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
por: Ackerman, Christopher, et al.
Publicado: (2024)
por: Ackerman, Christopher, et al.
Publicado: (2024)
TracrBench: Generating Interpretability Testbeds with Large Language Models
por: Thurnherr, Hannes, et al.
Publicado: (2024)
por: Thurnherr, Hannes, et al.
Publicado: (2024)
Mitigating Many-Shot Jailbreaking
por: Ackerman, Christopher M., et al.
Publicado: (2025)
por: Ackerman, Christopher M., et al.
Publicado: (2025)
Memory-Augmented Agent Training for Business Document Understanding
por: Liu, Jiale, et al.
Publicado: (2024)
por: Liu, Jiale, et al.
Publicado: (2024)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
por: Ball, Sarah, et al.
Publicado: (2024)
por: Ball, Sarah, et al.
Publicado: (2024)
Unlocking Cognitive Capabilities and Analyzing the Perception-Logic Trade-off
por: Zhang, Longyin, et al.
Publicado: (2026)
por: Zhang, Longyin, et al.
Publicado: (2026)
Towards evaluations-based safety cases for AI scheming
por: Balesni, Mikita, et al.
Publicado: (2024)
por: Balesni, Mikita, et al.
Publicado: (2024)
Think Smart, Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning
por: Chatterji, Satchit, et al.
Publicado: (2024)
por: Chatterji, Satchit, et al.
Publicado: (2024)
Discovering and Learning Probabilistic Models of Black-Box AI Capabilities
por: Bramblett, Daniel, et al.
Publicado: (2025)
por: Bramblett, Daniel, et al.
Publicado: (2025)
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
por: Li, Xiangyi, et al.
Publicado: (2026)
por: Li, Xiangyi, et al.
Publicado: (2026)
REBUS: A Robust Evaluation Benchmark of Understanding Symbols
por: Gritsevskiy, Andrew, et al.
Publicado: (2024)
por: Gritsevskiy, Andrew, et al.
Publicado: (2024)
Will we run out of data? Limits of LLM scaling based on human-generated data
por: Villalobos, Pablo, et al.
Publicado: (2022)
por: Villalobos, Pablo, et al.
Publicado: (2022)
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
por: Fu, Lingyue, et al.
Publicado: (2025)
por: Fu, Lingyue, et al.
Publicado: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
por: Black, Sid, et al.
Publicado: (2025)
por: Black, Sid, et al.
Publicado: (2025)
Log analysis is necessary for credible evaluation of AI agents
por: Kirgis, Peter, et al.
Publicado: (2026)
por: Kirgis, Peter, et al.
Publicado: (2026)
Intentional Deception as Controllable Capability in LLM Agents
por: Starace, Jason, et al.
Publicado: (2026)
por: Starace, Jason, et al.
Publicado: (2026)
Evaluating Environments Using Exploratory Agents
por: Khaleque, Bobby, et al.
Publicado: (2024)
por: Khaleque, Bobby, et al.
Publicado: (2024)
Evaluating Developmental Cognition Capabilities of LLMs
por: Xiao, Xiao, et al.
Publicado: (2026)
por: Xiao, Xiao, et al.
Publicado: (2026)
Tracking Capabilities for Safer Agents
por: Odersky, Martin, et al.
Publicado: (2026)
por: Odersky, Martin, et al.
Publicado: (2026)
Analyzing and Internalizing Complex Policy Documents for LLM Agents
por: Liu, Jiateng, et al.
Publicado: (2025)
por: Liu, Jiateng, et al.
Publicado: (2025)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
por: Olson, Matthew Lyle, et al.
Publicado: (2026)
por: Olson, Matthew Lyle, et al.
Publicado: (2026)
Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation
por: Sudhakar, Arjun Vaithilingam
Publicado: (2025)
por: Sudhakar, Arjun Vaithilingam
Publicado: (2025)
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
por: Backlund, Axel, et al.
Publicado: (2025)
por: Backlund, Axel, et al.
Publicado: (2025)
A Conceptual Framework for AI Capability Evaluations
por: Carro, María Victoria, et al.
Publicado: (2025)
por: Carro, María Victoria, et al.
Publicado: (2025)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
por: Turk, Matt
Publicado: (2026)
por: Turk, Matt
Publicado: (2026)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
por: Anurin, Andrey, et al.
Publicado: (2024)
por: Anurin, Andrey, et al.
Publicado: (2024)
Operand Quant: A Single-Agent Architecture for Autonomous Machine Learning Engineering
por: Sahney, Arjun, et al.
Publicado: (2025)
por: Sahney, Arjun, et al.
Publicado: (2025)
CADMAS-CTX: Contextual Capability Calibration for Multi-Agent Delegation
por: Qiao, Chuhan
Publicado: (2026)
por: Qiao, Chuhan
Publicado: (2026)
Responsibility-aware Strategic Reasoning in Probabilistic Multi-Agent Systems
por: Mu, Chunyan, et al.
Publicado: (2024)
por: Mu, Chunyan, et al.
Publicado: (2024)
Ejemplares similares
-
Forecasting Frontier Language Model Agent Capabilities
por: Pimpale, Govind, et al.
Publicado: (2025) -
Large Language Models Often Know When They Are Being Evaluated
por: Needham, Joe, et al.
Publicado: (2025) -
Large Language Models can Strategically Deceive their Users when Put Under Pressure
por: Scheurer, Jérémy, et al.
Publicado: (2023) -
Frontier Models are Capable of In-context Scheming
por: Meinke, Alexander, et al.
Publicado: (2024) -
Applying Refusal-Vector Ablation to Llama 3.1 70B Agents
por: Lermen, Simon, et al.
Publicado: (2024)