Log analysis is necessary for credible evaluation of AI agents
Fuente:
arXiv
Saved in:
| Main Authors: | Kirgis, Peter, Kapoor, Sayash, Rabanser, Stephan, Nadgir, Nitya, Ududec, Cozmin, Dubois, Magda, Allaire, JJ, Stosz, Conrad, Hobbhahn, Marius, Steinhardt, Jacob, Narayanan, Arvind |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards a Science of AI Agent Reliability
by: Rabanser, Stephan, et al.
Published: (2026)
by: Rabanser, Stephan, et al.
Published: (2026)
AI Agents That Matter
by: Kapoor, Sayash, et al.
Published: (2024)
by: Kapoor, Sayash, et al.
Published: (2024)
Open-World Evaluations for Measuring Frontier AI Capabilities
by: Kapoor, Sayash, et al.
Published: (2026)
by: Kapoor, Sayash, et al.
Published: (2026)
Ask don't tell: Reducing sycophancy in large language models
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
by: Siegel, Zachary S., et al.
Published: (2024)
by: Siegel, Zachary S., et al.
Published: (2024)
The Limits of Inference Scaling Through Resampling
by: Stroebl, Benedikt, et al.
Published: (2024)
by: Stroebl, Benedikt, et al.
Published: (2024)
Promises and pitfalls of artificial intelligence for legal applications
by: Kapoor, Sayash, et al.
Published: (2024)
by: Kapoor, Sayash, et al.
Published: (2024)
Skewed Score: A statistical framework to assess autograders
by: Dubois, Magda, et al.
Published: (2025)
by: Dubois, Magda, et al.
Published: (2025)
Seven simple steps for log analysis in AI systems
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning
by: Rabanser, Stephan
Published: (2025)
by: Rabanser, Stephan
Published: (2025)
What Does It Take to Build a Performant Selective Classifier?
by: Rabanser, Stephan, et al.
Published: (2025)
by: Rabanser, Stephan, et al.
Published: (2025)
Foundation Model Transparency Reports
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
Differences in the Moral Foundations of Large Language Models
by: Kirgis, Peter
Published: (2025)
by: Kirgis, Peter
Published: (2025)
Build Agent Advocates, Not Platform Agents
by: Kapoor, Sayash, et al.
Published: (2025)
by: Kapoor, Sayash, et al.
Published: (2025)
Quantifying Media Influence on Covid-19 Mask-Wearing Beliefs
by: Rabb, Nicholas, et al.
Published: (2024)
by: Rabb, Nicholas, et al.
Published: (2024)
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
by: Summerfield, Christopher, et al.
Published: (2025)
by: Summerfield, Christopher, et al.
Published: (2025)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
by: Kapoor, Sayash, et al.
Published: (2025)
by: Kapoor, Sayash, et al.
Published: (2025)
Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings
by: Pouget, Angéline, et al.
Published: (2025)
by: Pouget, Angéline, et al.
Published: (2025)
Technical Report: Evaluating Goal Drift in Language Model Agents
by: Arike, Rauno, et al.
Published: (2025)
by: Arike, Rauno, et al.
Published: (2025)
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
Establishing Best Practices for Building Rigorous Agentic Benchmarks
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Detecting Strategic Deception Using Linear Probes
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025)
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025)
How does Chain of Thought decompose complex tasks?
by: Nadgir, Amrut, et al.
Published: (2026)
by: Nadgir, Amrut, et al.
Published: (2026)
Context-Masked Meta-Prompting for Privacy-Preserving LLM Adaptation in Finance
by: Hiraou, Sayash Raaj
Published: (2024)
by: Hiraou, Sayash Raaj
Published: (2024)
Large Language Models Often Know When They Are Being Evaluated
by: Needham, Joe, et al.
Published: (2025)
by: Needham, Joe, et al.
Published: (2025)
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
by: Højmark, Axel, et al.
Published: (2024)
by: Højmark, Axel, et al.
Published: (2024)
Security policy audits: why and how
by: Narayanan, Arvind, et al.
Published: (2022)
by: Narayanan, Arvind, et al.
Published: (2022)
On the scientific credibility of paleoanthropology
by: Brian Villmoare, et al.
Published: (2024)
by: Brian Villmoare, et al.
Published: (2024)
Shaved Ice: Optimal Compute Resource Commitments for Dynamic Multi-Cloud Workloads
by: Stokely, Murray, et al.
Published: (2025)
by: Stokely, Murray, et al.
Published: (2025)
The 2024 Foundation Model Transparency Index
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024)
by: Meinke, Alexander, et al.
Published: (2024)
Will we run out of data? Limits of LLM scaling based on human-generated data
by: Villalobos, Pablo, et al.
Published: (2022)
by: Villalobos, Pablo, et al.
Published: (2022)
Flexible inference in heterogeneous and attributed multilayer networks
by: Contisciani, Martina, et al.
Published: (2024)
by: Contisciani, Martina, et al.
Published: (2024)
Students' credibility criteria for evaluating scientific information: The case of climate change on social media
by: Soraya Kresin, et al.
Published: (2024)
by: Soraya Kresin, et al.
Published: (2024)
Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
by: Rabanser, Stephan, et al.
Published: (2025)
by: Rabanser, Stephan, et al.
Published: (2025)
AutoFreeFem: Automatic code generation with FreeFEM++ and LaTex output for shape and topology optimization of non-linear multi-physics problems
by: Allaire, Grégoire, et al.
Published: (2024)
by: Allaire, Grégoire, et al.
Published: (2024)
Children's anthropomorphism of inanimate agents
by: Elizabeth J. Goldman, et al.
Published: (2024)
by: Elizabeth J. Goldman, et al.
Published: (2024)
Inflation target credibility and the Taylor rule
by: Dieter Nautz
Published: (2026)
by: Dieter Nautz
Published: (2026)
Similar Items
-
Towards a Science of AI Agent Reliability
by: Rabanser, Stephan, et al.
Published: (2026) -
AI Agents That Matter
by: Kapoor, Sayash, et al.
Published: (2024) -
Open-World Evaluations for Measuring Frontier AI Capabilities
by: Kapoor, Sayash, et al.
Published: (2026) -
Ask don't tell: Reducing sycophancy in large language models
by: Dubois, Magda, et al.
Published: (2026) -
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025)