Rigorous Interpretation Is a Form of Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Isabelle, Liu, Emmy, Jiao, Cathy, Joshi, Brihi, Yogatama, Dani, Barez, Fazl, Saxon, Michael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking AI Cultural Alignment
por: Bravansky, Michal, et al.
Publicado: (2025)
por: Bravansky, Michal, et al.
Publicado: (2025)
Token Taxes: mitigating AGI's economic risks
por: Irwin, Lucas, et al.
Publicado: (2026)
por: Irwin, Lucas, et al.
Publicado: (2026)
Evaluating Large Language Models for Fair and Reliable Organ Allocation
por: Kim, Brian Hyeongseok, et al.
Publicado: (2025)
por: Kim, Brian Hyeongseok, et al.
Publicado: (2025)
Embodied AI: Emerging Risks and Opportunities for Policy Action
por: Perlo, Jared, et al.
Publicado: (2025)
por: Perlo, Jared, et al.
Publicado: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
por: Lan, Michael, et al.
Publicado: (2023)
por: Lan, Michael, et al.
Publicado: (2023)
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
por: Lee, Isabelle, et al.
Publicado: (2024)
por: Lee, Isabelle, et al.
Publicado: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
por: Chaudhary, Maheep, et al.
Publicado: (2025)
por: Chaudhary, Maheep, et al.
Publicado: (2025)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
por: Fu, Tingchen, et al.
Publicado: (2025)
por: Fu, Tingchen, et al.
Publicado: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
por: Neo, Clement, et al.
Publicado: (2024)
por: Neo, Clement, et al.
Publicado: (2024)
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale
por: Lee, Isabelle, et al.
Publicado: (2025)
por: Lee, Isabelle, et al.
Publicado: (2025)
OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants
por: Ranjit, Jaspreet, et al.
Publicado: (2024)
por: Ranjit, Jaspreet, et al.
Publicado: (2024)
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
por: Quirke, Philip, et al.
Publicado: (2025)
por: Quirke, Philip, et al.
Publicado: (2025)
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
por: Chiang, Ting-Rui, et al.
Publicado: (2025)
por: Chiang, Ting-Rui, et al.
Publicado: (2025)
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
por: Chiang, Ting-Rui, et al.
Publicado: (2024)
por: Chiang, Ting-Rui, et al.
Publicado: (2024)
Understanding Addition and Subtraction in Transformers
por: Quirke, Philip, et al.
Publicado: (2024)
por: Quirke, Philip, et al.
Publicado: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
por: Gupta, Aman, et al.
Publicado: (2025)
por: Gupta, Aman, et al.
Publicado: (2025)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
por: Olteanu, Alexandra, et al.
Publicado: (2025)
por: Olteanu, Alexandra, et al.
Publicado: (2025)
Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
por: Wei, Kevin L., et al.
Publicado: (2025)
por: Wei, Kevin L., et al.
Publicado: (2025)
A Taxonomy of Response Strategies to Toxic Online Content: Evaluating the Evidence
por: Schirch, Lisa, et al.
Publicado: (2025)
por: Schirch, Lisa, et al.
Publicado: (2025)
Understanding Addition in Transformers
por: Quirke, Philip, et al.
Publicado: (2023)
por: Quirke, Philip, et al.
Publicado: (2023)
On Retrieval Augmentation and the Limitations of Language Model Training
por: Chiang, Ting-Rui, et al.
Publicado: (2023)
por: Chiang, Ting-Rui, et al.
Publicado: (2023)
Towards Interpreting Visual Information Processing in Vision-Language Models
por: Neo, Clement, et al.
Publicado: (2024)
por: Neo, Clement, et al.
Publicado: (2024)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
por: Liu, Ollie, et al.
Publicado: (2024)
por: Liu, Ollie, et al.
Publicado: (2024)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
por: Marks, Luke, et al.
Publicado: (2024)
por: Marks, Luke, et al.
Publicado: (2024)
LocateBench: Evaluating the Locating Ability of Vision Language Models
por: Chiang, Ting-Rui, et al.
Publicado: (2024)
por: Chiang, Ting-Rui, et al.
Publicado: (2024)
Query Circuits: Explaining How Language Models Answer User Prompts
por: Wu, Tung-Yu, et al.
Publicado: (2025)
por: Wu, Tung-Yu, et al.
Publicado: (2025)
Black-Box Access is Insufficient for Rigorous AI Audits
por: Casper, Stephen, et al.
Publicado: (2024)
por: Casper, Stephen, et al.
Publicado: (2024)
Smitten: A Trainee's Love of ACE Units
por: Emmy Yang
Publicado: (2025)
por: Emmy Yang
Publicado: (2025)
PrimeX: A Dataset of Worldview, Opinion, and Explanation
por: Koncel-Kedziorski, Rik, et al.
Publicado: (2025)
por: Koncel-Kedziorski, Rik, et al.
Publicado: (2025)
What is a typical signalized intersection in a city? A pipeline for intersection data imputation from OpenStreetMap
por: Qu, Ao, et al.
Publicado: (2024)
por: Qu, Ao, et al.
Publicado: (2024)
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
por: Bucknall, Ben, et al.
Publicado: (2025)
por: Bucknall, Ben, et al.
Publicado: (2025)
ExOAR: Expert-Guided Object and Activity Recognition from Textual Data
por: Beerepoot, Iris, et al.
Publicado: (2025)
por: Beerepoot, Iris, et al.
Publicado: (2025)
Integrating Wearable Data into Process Mining: Event, Case and Activity Enrichment
por: Dani, Vinicius Stein, et al.
Publicado: (2025)
por: Dani, Vinicius Stein, et al.
Publicado: (2025)
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
por: Brundage, Miles, et al.
Publicado: (2026)
por: Brundage, Miles, et al.
Publicado: (2026)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
por: Saxon, Michael, et al.
Publicado: (2024)
por: Saxon, Michael, et al.
Publicado: (2024)
Transparent AI: The Case for Interpretability and Explainability
por: Ramachandram, Dhanesh, et al.
Publicado: (2025)
por: Ramachandram, Dhanesh, et al.
Publicado: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
por: Simhi, Adi, et al.
Publicado: (2025)
por: Simhi, Adi, et al.
Publicado: (2025)
Mind the Gap! Pathways Towards Unifying AI Safety and Ethics Research
por: Roytburg, Dani, et al.
Publicado: (2025)
por: Roytburg, Dani, et al.
Publicado: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
por: Oozeer, Narmeen, et al.
Publicado: (2025)
por: Oozeer, Narmeen, et al.
Publicado: (2025)
Scaling sparse feature circuit finding for in-context learning
por: Kharlapenko, Dmitrii, et al.
Publicado: (2025)
por: Kharlapenko, Dmitrii, et al.
Publicado: (2025)
Ejemplares similares
-
Rethinking AI Cultural Alignment
por: Bravansky, Michal, et al.
Publicado: (2025) -
Token Taxes: mitigating AGI's economic risks
por: Irwin, Lucas, et al.
Publicado: (2026) -
Evaluating Large Language Models for Fair and Reliable Organ Allocation
por: Kim, Brian Hyeongseok, et al.
Publicado: (2025) -
Embodied AI: Emerging Risks and Opportunities for Policy Action
por: Perlo, Jared, et al.
Publicado: (2025) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
por: Lan, Michael, et al.
Publicado: (2023)