Establishing Best Practices for Building Rigorous Agentic Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Yuxuan, Jin, Tengjun, Pruksachatkun, Yada, Zhang, Andy, Liu, Shu, Cui, Sasha, Kapoor, Sayash, Longpre, Shayne, Meng, Kevin, Weiss, Rebecca, Barez, Fazl, Gupta, Rahul, Dhamala, Jwala, Merizian, Jacob, Giulianelli, Mario, Coppock, Harry, Ududec, Cozmin, Sekhon, Jasjeet, Steinhardt, Jacob, Kellermann, Antony, Schwettmann, Sarah, Zaharia, Matei, Stoica, Ion, Liang, Percy, Kang, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Skewed Score: A statistical framework to assess autograders
by: Dubois, Magda, et al.
Published: (2025)
by: Dubois, Magda, et al.
Published: (2025)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
Log analysis is necessary for credible evaluation of AI agents
by: Kirgis, Peter, et al.
Published: (2026)
by: Kirgis, Peter, et al.
Published: (2026)
ADAG: Automatically Describing Attribution Graphs
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
Language Model Circuits Are Sparse in the Neuron Basis
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
The 2024 Foundation Model Transparency Index
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
Foundation Model Transparency Reports
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
The 2025 Foundation Model Transparency Index
by: Wan, Alexander, et al.
Published: (2025)
by: Wan, Alexander, et al.
Published: (2025)
Ask don't tell: Reducing sycophancy in large language models
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023)
by: Quirke, Philip, et al.
Published: (2023)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
Future and AI-Ready Data Strategies: Response to DOC RFI on AI and Open Government Data Assets
by: Oderinwale, Hamidah, et al.
Published: (2024)
by: Oderinwale, Hamidah, et al.
Published: (2024)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
by: Huang, Vincent, et al.
Published: (2025)
by: Huang, Vincent, et al.
Published: (2025)
DECOR: Auditing LLM Deception via Information Manipulation Theory
by: Cai, Linyue, et al.
Published: (2026)
by: Cai, Linyue, et al.
Published: (2026)
Query Circuits: Explaining How Language Models Answer User Prompts
by: Wu, Tung-Yu, et al.
Published: (2025)
by: Wu, Tung-Yu, et al.
Published: (2025)
Algebraic and Statistical Properties of the Ordinary Least Squares Interpolator
by: Shen, Dennis, et al.
Published: (2023)
by: Shen, Dennis, et al.
Published: (2023)
Rethinking AI Cultural Alignment
by: Bravansky, Michal, et al.
Published: (2025)
by: Bravansky, Michal, et al.
Published: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
Long Context RAG Performance of Large Language Models
by: Leng, Quinn, et al.
Published: (2024)
by: Leng, Quinn, et al.
Published: (2024)
Token Taxes: mitigating AGI's economic risks
by: Irwin, Lucas, et al.
Published: (2026)
by: Irwin, Lucas, et al.
Published: (2026)
Large Language Models Relearn Removed Concepts
by: Lo, Michelle, et al.
Published: (2024)
by: Lo, Michelle, et al.
Published: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Time-uniform Chernoff bounds via nonnegative supermartingales
by: Howard, Steven R., et al.
Published: (2018)
by: Howard, Steven R., et al.
Published: (2018)
Open-World Evaluations for Measuring Frontier AI Capabilities
by: Kapoor, Sayash, et al.
Published: (2026)
by: Kapoor, Sayash, et al.
Published: (2026)
Seven simple steps for log analysis in AI systems
by: Dubois, Magda, et al.
Published: (2026)
by: Dubois, Magda, et al.
Published: (2026)
InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
The Need for a Socially-Grounded Persona Framework for User Simulation
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
How do Language Models Bind Entities in Context?
by: Feng, Jiahai, et al.
Published: (2023)
by: Feng, Jiahai, et al.
Published: (2023)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
The Limits of Inference Scaling Through Resampling
by: Stroebl, Benedikt, et al.
Published: (2024)
by: Stroebl, Benedikt, et al.
Published: (2024)
Build Agent Advocates, Not Platform Agents
by: Kapoor, Sayash, et al.
Published: (2025)
by: Kapoor, Sayash, et al.
Published: (2025)
Promises and pitfalls of artificial intelligence for legal applications
by: Kapoor, Sayash, et al.
Published: (2024)
by: Kapoor, Sayash, et al.
Published: (2024)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026)
by: Chen, Lingjiao, et al.
Published: (2026)
LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024)
by: Marks, Luke, et al.
Published: (2024)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
by: Heindrich, Lovis, et al.
Published: (2025)
by: Heindrich, Lovis, et al.
Published: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
Similar Items
-
Skewed Score: A statistical framework to assess autograders
by: Dubois, Magda, et al.
Published: (2025) -
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025) -
Log analysis is necessary for credible evaluation of AI agents
by: Kirgis, Peter, et al.
Published: (2026) -
ADAG: Automatically Describing Attribution Graphs
by: Arora, Aryaman, et al.
Published: (2026) -
Language Model Circuits Are Sparse in the Neuron Basis
by: Arora, Aryaman, et al.
Published: (2026)