The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Richard, Agarwal, Arunim, Mazeika, Mantas, Menghini, Cristina, Vacareanu, Robert, Kenstler, Brad, Yang, Mick, Barrass, Isabelle, Gatti, Alice, Yin, Xuwang, Trevino, Eduardo, Geralnik, Matias, Khoja, Adam, Lee, Dean, Yue, Summer, Hendrycks, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
TextQuests: How Good are LLMs at Text-Based Video Games?
by: Phan, Long, et al.
Published: (2025)
by: Phan, Long, et al.
Published: (2025)
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
by: Mazeika, Mantas, et al.
Published: (2025)
by: Mazeika, Mantas, et al.
Published: (2025)
Aggressive Compression Enables LLM Weight Theft
by: Brown, Davis, et al.
Published: (2026)
by: Brown, Davis, et al.
Published: (2026)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
by: Wang, Clinton J., et al.
Published: (2025)
by: Wang, Clinton J., et al.
Published: (2025)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
by: Mazeika, Mantas, et al.
Published: (2024)
by: Mazeika, Mantas, et al.
Published: (2024)
Reducing Political Manipulation with Consistency Training
by: Phan, Long, et al.
Published: (2026)
by: Phan, Long, et al.
Published: (2026)
Remote Labor Index: Measuring AI Automation of Remote Work
by: Mazeika, Mantas, et al.
Published: (2025)
by: Mazeika, Mantas, et al.
Published: (2025)
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
by: Wei, Boyi, et al.
Published: (2025)
by: Wei, Boyi, et al.
Published: (2025)
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification
by: Chen, Claire, et al.
Published: (2026)
by: Chen, Claire, et al.
Published: (2026)
Accuracy of Mathematical Functions in Julia
by: Mikaitis, Mantas, et al.
Published: (2025)
by: Mikaitis, Mantas, et al.
Published: (2025)
Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
by: Lauffer, Niklas, et al.
Published: (2025)
by: Lauffer, Niklas, et al.
Published: (2025)
Accuracy Is Not Honesty in AI Models: Human Agency and the Dynamics of Interaction
by: Matta, David (Daoud)
Published: (2026)
by: Matta, David (Daoud)
Published: (2026)
Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
by: Mohamadi, Mohamad Amin, et al.
Published: (2025)
by: Mohamadi, Mohamad Amin, et al.
Published: (2025)
MAGIC-MASK: Multi-Agent Guided Inter-Agent Collaboration with Mask-Based Explainability for Reinforcement Learning
by: Maliha, Maisha, et al.
Published: (2025)
by: Maliha, Maisha, et al.
Published: (2025)
Alignment for Honesty
by: Yang, Yuqing, et al.
Published: (2023)
by: Yang, Yuqing, et al.
Published: (2023)
Private equity acquisitions and product market decisions: Evidence from trademarks
by: Moazzam Khoja
Published: (2025)
by: Moazzam Khoja
Published: (2025)
Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privacy Policies
by: Ali, Mir Masood, et al.
Published: (2023)
by: Ali, Mir Masood, et al.
Published: (2023)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Leveraging Quantum Computing for Accelerated Classical Algorithms in Power Systems Optimization
by: Barrass, Rosemary, et al.
Published: (2025)
by: Barrass, Rosemary, et al.
Published: (2025)
[MASK] is All You Need
by: Hu, Vincent Tao, et al.
Published: (2024)
by: Hu, Vincent Tao, et al.
Published: (2024)
Pini, M., Más Rocha, S., Gorostiaga, J., Tello, C., y Asprella, G. (coords.) La Educación Secundaria, ¿Modelo en (re)construcción?. Buenos Aires: Editorial Aique, pags. 252.
by: Raúl A. Menghini
Published: (2016)
by: Raúl A. Menghini
Published: (2016)
Presentación
by: Raúl A. Menghini
Published: (2022)
by: Raúl A. Menghini
Published: (2022)
Forging the Ideal Educated Girl (Volume 1.0)
by: Khoja-Moolji, Shenila
Published: (2020)
by: Khoja-Moolji, Shenila
Published: (2020)
Forging the Ideal Educated Girl
by: Khoja-Moolji, Shenila
Published: (2018)
by: Khoja-Moolji, Shenila
Published: (2018)
Representation Engineering: A Top-Down Approach to AI Transparency
by: Zou, Andy, et al.
Published: (2023)
by: Zou, Andy, et al.
Published: (2023)
Training LLMs for Honesty via Confessions
by: Joglekar, Manas, et al.
Published: (2025)
by: Joglekar, Manas, et al.
Published: (2025)
Annotation-Efficient Universal Honesty Alignment
by: Ni, Shiyu, et al.
Published: (2025)
by: Ni, Shiyu, et al.
Published: (2025)
Conditional [MASK] Discrete Diffusion Language Model
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Two‐stage pricing of products and services considering different competitive environments
by: Wei Qi, et al.
Published: (2024)
by: Wei Qi, et al.
Published: (2024)
A Survey on the Honesty of Large Language Models
by: Li, Siheng, et al.
Published: (2024)
by: Li, Siheng, et al.
Published: (2024)
Honesty, Stigma, and Cooperation in an Overlapping-Generations Game
by: Li, David, et al.
Published: (2025)
by: Li, David, et al.
Published: (2025)
Energy Transformation in Lithuania
by: Švažas, Mantas
Published: (2025)
by: Švažas, Mantas
Published: (2025)
Scaphoid Fracture Reconstruction with Rib Autograft: Case Report and Literature Review
by: Mantas Fomkinas
Published: (2021)
by: Mantas Fomkinas
Published: (2021)
What’s so Funny? Democritus ridens in Juvenal 10
by: Mantas Adomėnas
Published: (2022)
by: Mantas Adomėnas
Published: (2022)
MATLAB Simulator of Level-Index Arithmetic
by: Mikaitis, Mantas
Published: (2024)
by: Mikaitis, Mantas
Published: (2024)
Istorinio pasakojimo prielaidos Augustino De civitate Dei
by: Mantas Tamošaitis
Published: (2021)
by: Mantas Tamošaitis
Published: (2021)
Incentivizing Honesty among Competitors in Collaborative Learning and Optimization
by: Dorner, Florian E., et al.
Published: (2023)
by: Dorner, Florian E., et al.
Published: (2023)
Similar Items
-
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024) -
TextQuests: How Good are LLMs at Text-Based Video Games?
by: Phan, Long, et al.
Published: (2025) -
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
by: Mazeika, Mantas, et al.
Published: (2025) -
Aggressive Compression Enables LLM Weight Theft
by: Brown, Davis, et al.
Published: (2026) -
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
by: Wang, Clinton J., et al.
Published: (2025)