Black-Box Access is Insufficient for Rigorous AI Audits
Fuente:
arXiv
Saved in:
| Main Authors: | Casper, Stephen, Ezell, Carson, Siegmann, Charlotte, Kolt, Noam, Curtis, Taylor Lynn, Bucknall, Benjamin, Haupt, Andreas, Wei, Kevin, Scheurer, Jérémy, Hobbhahn, Marius, Sharkey, Lee, Krishna, Satyapriya, Von Hagen, Marvin, Alberti, Silas, Chan, Alan, Sun, Qinyi, Gerovitch, Michael, Bau, David, Tegmark, Max, Krueger, David, Hadfield-Menell, Dylan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pitfalls of Evidence-Based AI Policy
by: Casper, Stephen, et al.
Published: (2025)
by: Casper, Stephen, et al.
Published: (2025)
The AI Agent Index
by: Casper, Stephen, et al.
Published: (2025)
by: Casper, Stephen, et al.
Published: (2025)
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
by: Khan, Ariba, et al.
Published: (2025)
by: Khan, Ariba, et al.
Published: (2025)
Governing AI Agents
by: Kolt, Noam
Published: (2025)
by: Kolt, Noam
Published: (2025)
Superintelligence and Law
by: Kolt, Noam
Published: (2026)
by: Kolt, Noam
Published: (2026)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
Visibility into AI Agents
by: Chan, Alan, et al.
Published: (2024)
by: Chan, Alan, et al.
Published: (2024)
Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL
by: Haupt, Andreas A., et al.
Published: (2022)
by: Haupt, Andreas A., et al.
Published: (2022)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Eight Methods to Evaluate Robust Unlearning in LLMs
by: Lynch, Aengus, et al.
Published: (2024)
by: Lynch, Aengus, et al.
Published: (2024)
An FDA for AI? Pitfalls and Plausibility of Approval Regulation for Frontier Artificial Intelligence
by: Carpenter, Daniel, et al.
Published: (2024)
by: Carpenter, Daniel, et al.
Published: (2024)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
by: Ma, Rachel, et al.
Published: (2026)
by: Ma, Rachel, et al.
Published: (2026)
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
by: Soni, Prajna, et al.
Published: (2025)
by: Soni, Prajna, et al.
Published: (2025)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
by: Hahm, Dongyoon, et al.
Published: (2026)
by: Hahm, Dongyoon, et al.
Published: (2026)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
by: Sankaranarayanan, Aruna, et al.
Published: (2025)
by: Sankaranarayanan, Aruna, et al.
Published: (2025)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
by: Siththaranjan, Anand, et al.
Published: (2023)
by: Siththaranjan, Anand, et al.
Published: (2023)
Prompt Injection as Role Confusion
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
Build Agent Advocates, Not Platform Agents
by: Kapoor, Sayash, et al.
Published: (2025)
by: Kapoor, Sayash, et al.
Published: (2025)
Regulating AI Agents
by: Gardhouse, Kathrin, et al.
Published: (2026)
by: Gardhouse, Kathrin, et al.
Published: (2026)
Let's talk about sex: Why reproductive systems matter for understanding algae
by: Stacy A. Krueger‐Hadfield
Published: (2024)
by: Stacy A. Krueger‐Hadfield
Published: (2024)
Diverse Preference Learning for Capabilities and Alignment
by: Slocum, Stewart, et al.
Published: (2025)
by: Slocum, Stewart, et al.
Published: (2025)
Lessons from complexity theory for AI governance
by: Kolt, Noam, et al.
Published: (2025)
by: Kolt, Noam, et al.
Published: (2025)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
Cooperative Inverse Reinforcement Learning
by: Hadfield-Menell, Dylan, et al.
Published: (2016)
by: Hadfield-Menell, Dylan, et al.
Published: (2016)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2025)
by: Ma, Rachel, et al.
Published: (2025)
Layered Unlearning for Adversarial Relearning
by: Qian, Timothy, et al.
Published: (2025)
by: Qian, Timothy, et al.
Published: (2025)
Goal Inference from Open-Ended Dialog
by: Ma, Rachel, et al.
Published: (2024)
by: Ma, Rachel, et al.
Published: (2024)
Activation Steering via Generative Causal Mediation
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
by: Højmark, Axel, et al.
Published: (2024)
by: Højmark, Axel, et al.
Published: (2024)
Incident Analysis for AI Agents
by: Ezell, Carson, et al.
Published: (2025)
by: Ezell, Carson, et al.
Published: (2025)
PRIMES STEP Experience
by: Gerovitch, Slava, et al.
Published: (2026)
by: Gerovitch, Slava, et al.
Published: (2026)
Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains
by: Khan, Emaan Bilal, et al.
Published: (2026)
by: Khan, Emaan Bilal, et al.
Published: (2026)
Monoicy, dioicy, and genetic structure in three species of Sheathia (Batrachospermales, Rhodophyta)
by: Shainker-Connelly, Sarah, et al.
Published: (2025)
by: Shainker-Connelly, Sarah, et al.
Published: (2025)
Dark Speculation: Combining Qualitative and Quantitative Understanding in Frontier AI Risk Analysis
by: Carpenter, Daniel, et al.
Published: (2025)
by: Carpenter, Daniel, et al.
Published: (2025)
How Do AI Companies "Fine-Tune" Policy? Examining Regulatory Capture in AI Governance
by: Wei, Kevin, et al.
Published: (2024)
by: Wei, Kevin, et al.
Published: (2024)
On the formal ribbon extension of a quasitriangular Hopf algebra
by: Kolt, Quinn T.
Published: (2024)
by: Kolt, Quinn T.
Published: (2024)
Climate of the Middle Understanding Climate Change as a Common Challenge
by: Arjen Siegmann
by: Arjen Siegmann
Caso-pensamento como estratégia na produção de conhecimento
by: Christiane Siegmann
Published: (2007)
by: Christiane Siegmann
Published: (2007)
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024)
by: Meinke, Alexander, et al.
Published: (2024)
Similar Items
-
Pitfalls of Evidence-Based AI Policy
by: Casper, Stephen, et al.
Published: (2025) -
The AI Agent Index
by: Casper, Stephen, et al.
Published: (2025) -
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
by: Khan, Ariba, et al.
Published: (2025) -
Governing AI Agents
by: Kolt, Noam
Published: (2025) -
Superintelligence and Law
by: Kolt, Noam
Published: (2026)