AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors
Fuente:
arXiv
Saved in:
| Main Authors: | Sheshadri, Abhay, Ewart, Aidan, Fronsdal, Kai, Gupta, Isha, Bowman, Samuel R., Price, Sara, Marks, Samuel, Wang, Rowan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024)
by: Guo, Phillip, et al.
Published: (2024)
Liars' Bench: Evaluating Lie Detectors for Language Models
by: Kretschmar, Kieron, et al.
Published: (2025)
by: Kretschmar, Kieron, et al.
Published: (2025)
DECOR: Auditing LLM Deception via Information Manipulation Theory
by: Cai, Linyue, et al.
Published: (2026)
by: Cai, Linyue, et al.
Published: (2026)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
by: Fronsdal, Kai, et al.
Published: (2024)
by: Fronsdal, Kai, et al.
Published: (2024)
Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
by: Pfau, Jacob, et al.
Published: (2024)
by: Pfau, Jacob, et al.
Published: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
Auditing language models for hidden objectives
by: Marks, Samuel, et al.
Published: (2025)
by: Marks, Samuel, et al.
Published: (2025)
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
by: Shenoy, Keshav, et al.
Published: (2026)
by: Shenoy, Keshav, et al.
Published: (2026)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026)
by: Tu, Xinming, et al.
Published: (2026)
AuditWen:An Open-Source Large Language Model for Audit
by: Huang, Jiajia, et al.
Published: (2024)
by: Huang, Jiajia, et al.
Published: (2024)
LLM Evaluators Recognize and Favor Their Own Generations
by: Panickssery, Arjun, et al.
Published: (2024)
by: Panickssery, Arjun, et al.
Published: (2024)
Audited Reasoning Refinement: Fine-Tuning Language Models via LLM-Guided Step-Wise Evaluation and Correction
by: Bhattacharyya, Sumanta, et al.
Published: (2025)
by: Bhattacharyya, Sumanta, et al.
Published: (2025)
What Artificial Neural Networks Can Tell Us About Human Language Acquisition
by: Warstadt, Alex, et al.
Published: (2022)
by: Warstadt, Alex, et al.
Published: (2022)
An Audit on the Perspectives and Challenges of Hallucinations in NLP
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
Market-Bench: Evaluating Large Language Models on Introductory Quantitative Trading and Market Dynamics
by: Srivastava, Abhay, et al.
Published: (2025)
by: Srivastava, Abhay, et al.
Published: (2025)
Audit, Alignment, and Optimization of LM-Powered Subroutines with Application to Public Comment Processing
by: Raab, Reilly, et al.
Published: (2025)
by: Raab, Reilly, et al.
Published: (2025)
Eight Methods to Evaluate Robust Unlearning in LLMs
by: Lynch, Aengus, et al.
Published: (2024)
by: Lynch, Aengus, et al.
Published: (2024)
GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence
by: Krishna, Kundan, et al.
Published: (2024)
by: Krishna, Kundan, et al.
Published: (2024)
AuditGPT: Auditing Smart Contracts with ChatGPT
by: Xia, Shihao, et al.
Published: (2024)
by: Xia, Shihao, et al.
Published: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
by: Patel, Bhrij, et al.
Published: (2024)
by: Patel, Bhrij, et al.
Published: (2024)
Model Spec Midtraining: Improving How Alignment Training Generalizes
by: Li, Chloe, et al.
Published: (2026)
by: Li, Chloe, et al.
Published: (2026)
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
by: Hua, Tim Tian, et al.
Published: (2025)
by: Hua, Tim Tian, et al.
Published: (2025)
Quantum-Audit: Evaluating the Reasoning Limits of LLMs on Quantum Computing
by: Afane, Mohamed, et al.
Published: (2026)
by: Afane, Mohamed, et al.
Published: (2026)
Auditing Rust Crates Effectively
by: Zoghbi, Lydia, et al.
Published: (2026)
by: Zoghbi, Lydia, et al.
Published: (2026)
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
by: Ridoy, Shahriyar Zaman, et al.
Published: (2025)
by: Ridoy, Shahriyar Zaman, et al.
Published: (2025)
Statistical Hypothesis Testing for Auditing Robustness in Language Models
by: Rauba, Paulius, et al.
Published: (2025)
by: Rauba, Paulius, et al.
Published: (2025)
Keeping Up with the Language Models: Systematic Benchmark Extension for Bias Auditing
by: Baldini, Ioana, et al.
Published: (2023)
by: Baldini, Ioana, et al.
Published: (2023)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning
by: Chen, Chaoran, et al.
Published: (2026)
by: Chen, Chaoran, et al.
Published: (2026)
Stage-Audit: Auditable Source-Frontier Discovery for Cross-Wiki Tables
by: Shen, Chen
Published: (2026)
by: Shen, Chen
Published: (2026)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
Auditing Agent Harness Safety
by: Liu, Chengzhi, et al.
Published: (2026)
by: Liu, Chengzhi, et al.
Published: (2026)
Smart Audit System Empowered by LLM
by: Yao, Xu, et al.
Published: (2024)
by: Yao, Xu, et al.
Published: (2024)
Direction-Flipped Influence Audits Reveal Hidden Structure in Moral Choices of LLMs
by: Blandfort, Phil, et al.
Published: (2026)
by: Blandfort, Phil, et al.
Published: (2026)
Auditing the Use of Language Models to Guide Hiring Decisions
by: Gaebler, Johann D., et al.
Published: (2024)
by: Gaebler, Johann D., et al.
Published: (2024)
Auditing Prompt Caching in Language Model APIs
by: Gu, Chenchen, et al.
Published: (2025)
by: Gu, Chenchen, et al.
Published: (2025)
Distilling Bayesian Belief States into Language Models for Auditable Negotiation
by: Cui, Zongqi, et al.
Published: (2026)
by: Cui, Zongqi, et al.
Published: (2026)
Automated Benchmark Auditing for AI Agents and Large Language Models
by: Wang, Junlin, et al.
Published: (2026)
by: Wang, Junlin, et al.
Published: (2026)
CALM: Curiosity-Driven Auditing for Large Language Models
by: Zheng, Xiang, et al.
Published: (2025)
by: Zheng, Xiang, et al.
Published: (2025)
Similar Items
-
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024) -
Liars' Bench: Evaluating Lie Detectors for Language Models
by: Kretschmar, Kieron, et al.
Published: (2025) -
DECOR: Auditing LLM Deception via Information Manipulation Theory
by: Cai, Linyue, et al.
Published: (2026) -
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
by: Fronsdal, Kai, et al.
Published: (2024) -
Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
by: Pfau, Jacob, et al.
Published: (2024)