Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Kevin L., Paskov, Patricia, Dev, Sunishchal, Byun, Michael J., Reuel, Anka, Roberts-Gaal, Xavier, Calcott, Rachel, Coxon, Evie, Deshpande, Chinmay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
by: Lamparth, Max, et al.
Published: (2023)
by: Lamparth, Max, et al.
Published: (2023)
Generative AI Needs Adaptive Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Fairness in Reinforcement Learning: A Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Preliminary suggestions for rigorous GPAI model evaluations
by: Paskov, Patricia, et al.
Published: (2025)
by: Paskov, Patricia, et al.
Published: (2025)
Audit Cards: Contextualizing AI Evaluations
by: Staufer, Leon, et al.
Published: (2025)
by: Staufer, Leon, et al.
Published: (2025)
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Incident Analysis for AI Agents
by: Ezell, Carson, et al.
Published: (2025)
by: Ezell, Carson, et al.
Published: (2025)
How Do AI Companies "Fine-Tune" Policy? Examining Regulatory Capture in AI Governance
by: Wei, Kevin, et al.
Published: (2024)
by: Wei, Kevin, et al.
Published: (2024)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
by: Haupt, Andreas, et al.
Published: (2026)
by: Haupt, Andreas, et al.
Published: (2026)
Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering
by: Mao, Nathan, et al.
Published: (2026)
by: Mao, Nathan, et al.
Published: (2026)
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
by: Rivera, Juan-Pablo, et al.
Published: (2024)
by: Rivera, Juan-Pablo, et al.
Published: (2024)
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
by: Sahay, Kenji, et al.
Published: (2025)
by: Sahay, Kenji, et al.
Published: (2025)
GPAI Evaluations Standards Taskforce: Towards Effective AI Governance
by: Paskov, Patricia, et al.
Published: (2024)
by: Paskov, Patricia, et al.
Published: (2024)
Evaluating LLMs in Medicine: A Call for Rigor, Transparency
by: Alwakeel, Mahmoud, et al.
Published: (2025)
by: Alwakeel, Mahmoud, et al.
Published: (2025)
Responsible AI in the Global Context: Maturity Model and Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
by: He, Jiahang, et al.
Published: (2025)
by: He, Jiahang, et al.
Published: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
An Adaptive Responsible AI Governance Framework for Decentralized Organizations
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
CA-BED: Conversation-Aware Bayesian Experimental Design
by: Arnould, Daniel, et al.
Published: (2026)
by: Arnould, Daniel, et al.
Published: (2026)
DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
by: Agrawal, Shriyansh, et al.
Published: (2025)
by: Agrawal, Shriyansh, et al.
Published: (2025)
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
by: Brundage, Miles, et al.
Published: (2026)
by: Brundage, Miles, et al.
Published: (2026)
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Parameter-Efficient Deep Learning for Ultrasound-Based Human-Machine Interfaces
by: Lykourinas, Antonios, et al.
Published: (2026)
by: Lykourinas, Antonios, et al.
Published: (2026)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025)
by: Batra, Shourya, et al.
Published: (2025)
Foundation Model Transparency Reports
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
Circle of Willis Centerline Graphs: A Dataset and Baseline Algorithm
by: Musio, Fabio, et al.
Published: (2025)
by: Musio, Fabio, et al.
Published: (2025)
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
by: Bucknall, Ben, et al.
Published: (2025)
by: Bucknall, Ben, et al.
Published: (2025)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
by: Afonin, Nikita, et al.
Published: (2025)
by: Afonin, Nikita, et al.
Published: (2025)
Call for Rigor in Reporting Quality of Instruction Tuning Data
by: Moon, Hyeonseok, et al.
Published: (2025)
by: Moon, Hyeonseok, et al.
Published: (2025)
Evaluating ChatGPT as a Recommender System: A Rigorous Approach
by: Di Palma, Dario, et al.
Published: (2023)
by: Di Palma, Dario, et al.
Published: (2023)
The science and practice of proportionality in AI risk evaluations
by: Mougan, Carlos, et al.
Published: (2026)
by: Mougan, Carlos, et al.
Published: (2026)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
LexChronos: An Agentic Framework for Structured Event Timeline Extraction in Indian Jurisprudence
by: Tummepalli, Anka Chandrahas, et al.
Published: (2026)
by: Tummepalli, Anka Chandrahas, et al.
Published: (2026)
GAMEBoT: Transparent Assessment of LLM Reasoning in Games
by: Lin, Wenye, et al.
Published: (2024)
by: Lin, Wenye, et al.
Published: (2024)
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
by: Zhou, Karen, et al.
Published: (2026)
by: Zhou, Karen, et al.
Published: (2026)
Large Language Models for Large-Scale, Rigorous Qualitative Analysis in Applied Health Services Research
by: Ronaghi, Sasha, et al.
Published: (2026)
by: Ronaghi, Sasha, et al.
Published: (2026)
Revisiting Simple Baselines for In-The-Wild Deepfake Detection
by: Castaneda, Orlando, et al.
Published: (2025)
by: Castaneda, Orlando, et al.
Published: (2025)
MiTTenS: A Dataset for Evaluating Gender Mistranslation
by: Robinson, Kevin, et al.
Published: (2024)
by: Robinson, Kevin, et al.
Published: (2024)
LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation
by: Byun, Grace, et al.
Published: (2025)
by: Byun, Grace, et al.
Published: (2025)
Similar Items
-
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
by: Lamparth, Max, et al.
Published: (2023) -
Generative AI Needs Adaptive Governance
by: Reuel, Anka, et al.
Published: (2024) -
Fairness in Reinforcement Learning: A Survey
by: Reuel, Anka, et al.
Published: (2024) -
Preliminary suggestions for rigorous GPAI model evaluations
by: Paskov, Patricia, et al.
Published: (2025) -
Audit Cards: Contextualizing AI Evaluations
by: Staufer, Leon, et al.
Published: (2025)