Audit Cards: Contextualizing AI Evaluations
Fuente:
arXiv
Saved in:
| Main Authors: | Staufer, Leon, Yang, Mick, Reuel, Anka, Casper, Stephen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative AI Needs Adaptive Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
by: Staufer, Leon, et al.
Published: (2026)
by: Staufer, Leon, et al.
Published: (2026)
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Fairness in Reinforcement Learning: A Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Responsible AI in the Global Context: Maturity Model and Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
What Do LLMs Associate with Your Name? A Human-Centered Black-Box Audit of Personal Data
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures
by: Stelling, Lily, et al.
Published: (2025)
by: Stelling, Lily, et al.
Published: (2025)
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services
by: Hartmann, David, et al.
Published: (2024)
by: Hartmann, David, et al.
Published: (2024)
Human-Centred LLM Privacy Audits: Findings and Frictions
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
Legal Alignment for Safe and Ethical AI
by: Kolt, Noam, et al.
Published: (2026)
by: Kolt, Noam, et al.
Published: (2026)
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
by: Rivera, Juan-Pablo, et al.
Published: (2024)
by: Rivera, Juan-Pablo, et al.
Published: (2024)
Expanding External Access To Frontier AI Models For Dangerous Capability Evaluations
by: Charnock, Jacob, et al.
Published: (2026)
by: Charnock, Jacob, et al.
Published: (2026)
Pitfalls of Evidence-Based AI Policy
by: Casper, Stephen, et al.
Published: (2025)
by: Casper, Stephen, et al.
Published: (2025)
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
by: Lamparth, Max, et al.
Published: (2023)
by: Lamparth, Max, et al.
Published: (2023)
An Adaptive Responsible AI Governance Framework for Decentralized Organizations
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
by: Brundage, Miles, et al.
Published: (2026)
by: Brundage, Miles, et al.
Published: (2026)
What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
by: Staufer, Dimitri
Published: (2025)
by: Staufer, Dimitri
Published: (2025)
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
by: Khan, Ariba, et al.
Published: (2025)
by: Khan, Ariba, et al.
Published: (2025)
Practical Principles for AI Cost and Compute Accounting
by: Casper, Stephen, et al.
Published: (2025)
by: Casper, Stephen, et al.
Published: (2025)
Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
by: Wei, Kevin L., et al.
Published: (2025)
by: Wei, Kevin L., et al.
Published: (2025)
AuditMAI: Towards An Infrastructure for Continuous AI Auditing
by: Waltersdorfer, Laura, et al.
Published: (2024)
by: Waltersdorfer, Laura, et al.
Published: (2024)
Deprecating Benchmarks: Criteria and Framework
by: Joaquin, Ayrton San, et al.
Published: (2025)
by: Joaquin, Ayrton San, et al.
Published: (2025)
Black-Box Access is Insufficient for Rigorous AI Audits
by: Casper, Stephen, et al.
Published: (2024)
by: Casper, Stephen, et al.
Published: (2024)
Impact Matters! An Audit Method to Evaluate AI Projects and their Impact for Sustainability and Public Interest
by: Züger, Theresa, et al.
Published: (2026)
by: Züger, Theresa, et al.
Published: (2026)
Open Problems in Technical AI Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Avoiding an AI-imposed Taylor's Version of all music history
by: Collins, Nick, et al.
Published: (2024)
by: Collins, Nick, et al.
Published: (2024)
Investigating Youth AI Auditing
by: Solyst, Jaemarie, et al.
Published: (2025)
by: Solyst, Jaemarie, et al.
Published: (2025)
Industrial AI Robustness Card for Time Series Models
by: Windmann, Alexander, et al.
Published: (2025)
by: Windmann, Alexander, et al.
Published: (2025)
Can AI be Auditable?
by: Verma, Himanshu, et al.
Published: (2025)
by: Verma, Himanshu, et al.
Published: (2025)
Synthesizing Proteins on the Graphics Card. Protein Folding and the Limits of Critical AI Studies
by: Offert, Fabian, et al.
Published: (2024)
by: Offert, Fabian, et al.
Published: (2024)
Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling
by: Ojewale, Victor, et al.
Published: (2024)
by: Ojewale, Victor, et al.
Published: (2024)
Evaluation Cards for XAI Metrics
by: Gipiškis, Rokas, et al.
Published: (2026)
by: Gipiškis, Rokas, et al.
Published: (2026)
EvalCards: A Framework for Standardized Evaluation Reporting
by: Dhar, Ruchira, et al.
Published: (2025)
by: Dhar, Ruchira, et al.
Published: (2025)
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
by: Gringras, David, et al.
Published: (2026)
by: Gringras, David, et al.
Published: (2026)
Evaluating Contextually Personalized Programming Exercises Created with Generative AI
by: Logacheva, Evanfiya, et al.
Published: (2024)
by: Logacheva, Evanfiya, et al.
Published: (2024)
A Blueprint for Auditing Generative AI
by: Mokander, Jakob, et al.
Published: (2024)
by: Mokander, Jakob, et al.
Published: (2024)
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
by: Bucknall, Ben, et al.
Published: (2025)
by: Bucknall, Ben, et al.
Published: (2025)
Human Experts' Evaluation of Generative AI for Contextualizing STEAM Education in the Global South
by: Nyaaba, Matthew, et al.
Published: (2025)
by: Nyaaba, Matthew, et al.
Published: (2025)
Similar Items
-
Generative AI Needs Adaptive Governance
by: Reuel, Anka, et al.
Published: (2024) -
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
by: Staufer, Leon, et al.
Published: (2026) -
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
by: Reuel, Anka, et al.
Published: (2024) -
Fairness in Reinforcement Learning: A Survey
by: Reuel, Anka, et al.
Published: (2024) -
Responsible AI in the Global Context: Maturity Model and Survey
by: Reuel, Anka, et al.
Published: (2024)