Rigorous Interpretation Is a Form of Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Isabelle, Liu, Emmy, Jiao, Cathy, Joshi, Brihi, Yogatama, Dani, Barez, Fazl, Saxon, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking AI Cultural Alignment
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
Token Taxes: mitigating AGI's economic risks
von: Irwin, Lucas, et al.
Veröffentlicht: (2026)
von: Irwin, Lucas, et al.
Veröffentlicht: (2026)
Evaluating Large Language Models for Fair and Reliable Organ Allocation
von: Kim, Brian Hyeongseok, et al.
Veröffentlicht: (2025)
von: Kim, Brian Hyeongseok, et al.
Veröffentlicht: (2025)
Embodied AI: Emerging Risks and Opportunities for Policy Action
von: Perlo, Jared, et al.
Veröffentlicht: (2025)
von: Perlo, Jared, et al.
Veröffentlicht: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
von: Lee, Isabelle, et al.
Veröffentlicht: (2024)
von: Lee, Isabelle, et al.
Veröffentlicht: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale
von: Lee, Isabelle, et al.
Veröffentlicht: (2025)
von: Lee, Isabelle, et al.
Veröffentlicht: (2025)
OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants
von: Ranjit, Jaspreet, et al.
Veröffentlicht: (2024)
von: Ranjit, Jaspreet, et al.
Veröffentlicht: (2024)
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
von: Quirke, Philip, et al.
Veröffentlicht: (2025)
von: Quirke, Philip, et al.
Veröffentlicht: (2025)
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2025)
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2025)
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2024)
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2024)
Understanding Addition and Subtraction in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
von: Olteanu, Alexandra, et al.
Veröffentlicht: (2025)
von: Olteanu, Alexandra, et al.
Veröffentlicht: (2025)
Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
von: Wei, Kevin L., et al.
Veröffentlicht: (2025)
von: Wei, Kevin L., et al.
Veröffentlicht: (2025)
A Taxonomy of Response Strategies to Toxic Online Content: Evaluating the Evidence
von: Schirch, Lisa, et al.
Veröffentlicht: (2025)
von: Schirch, Lisa, et al.
Veröffentlicht: (2025)
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
On Retrieval Augmentation and the Limitations of Language Model Training
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2023)
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2023)
Towards Interpreting Visual Information Processing in Vision-Language Models
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)
von: Marks, Luke, et al.
Veröffentlicht: (2024)
LocateBench: Evaluating the Locating Ability of Vision Language Models
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2024)
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2024)
Query Circuits: Explaining How Language Models Answer User Prompts
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
Black-Box Access is Insufficient for Rigorous AI Audits
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
Smitten: A Trainee's Love of ACE Units
von: Emmy Yang
Veröffentlicht: (2025)
von: Emmy Yang
Veröffentlicht: (2025)
PrimeX: A Dataset of Worldview, Opinion, and Explanation
von: Koncel-Kedziorski, Rik, et al.
Veröffentlicht: (2025)
von: Koncel-Kedziorski, Rik, et al.
Veröffentlicht: (2025)
What is a typical signalized intersection in a city? A pipeline for intersection data imputation from OpenStreetMap
von: Qu, Ao, et al.
Veröffentlicht: (2024)
von: Qu, Ao, et al.
Veröffentlicht: (2024)
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
von: Bucknall, Ben, et al.
Veröffentlicht: (2025)
von: Bucknall, Ben, et al.
Veröffentlicht: (2025)
ExOAR: Expert-Guided Object and Activity Recognition from Textual Data
von: Beerepoot, Iris, et al.
Veröffentlicht: (2025)
von: Beerepoot, Iris, et al.
Veröffentlicht: (2025)
Integrating Wearable Data into Process Mining: Event, Case and Activity Enrichment
von: Dani, Vinicius Stein, et al.
Veröffentlicht: (2025)
von: Dani, Vinicius Stein, et al.
Veröffentlicht: (2025)
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
von: Brundage, Miles, et al.
Veröffentlicht: (2026)
von: Brundage, Miles, et al.
Veröffentlicht: (2026)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
von: Saxon, Michael, et al.
Veröffentlicht: (2024)
Transparent AI: The Case for Interpretability and Explainability
von: Ramachandram, Dhanesh, et al.
Veröffentlicht: (2025)
von: Ramachandram, Dhanesh, et al.
Veröffentlicht: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Mind the Gap! Pathways Towards Unifying AI Safety and Ethics Research
von: Roytburg, Dani, et al.
Veröffentlicht: (2025)
von: Roytburg, Dani, et al.
Veröffentlicht: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rethinking AI Cultural Alignment
von: Bravansky, Michal, et al.
Veröffentlicht: (2025) -
Token Taxes: mitigating AGI's economic risks
von: Irwin, Lucas, et al.
Veröffentlicht: (2026) -
Evaluating Large Language Models for Fair and Reliable Organ Allocation
von: Kim, Brian Hyeongseok, et al.
Veröffentlicht: (2025) -
Embodied AI: Emerging Risks and Opportunities for Policy Action
von: Perlo, Jared, et al.
Veröffentlicht: (2025) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)