Pitfalls in Evaluating Interpretability Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Haklay, Tal, Prakash, Nikhil, Pandey, Sana, Torralba, Antonio, Mueller, Aaron, Andreas, Jacob, Shaham, Tamar Rott, Belinkov, Yonatan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Position-aware Automatic Circuit Discovery
por: Haklay, Tal, et al.
Publicado: (2025)
por: Haklay, Tal, et al.
Publicado: (2025)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
por: Arad, Dana, et al.
Publicado: (2023)
por: Arad, Dana, et al.
Publicado: (2023)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
por: Nikankin, Yaniv, et al.
Publicado: (2024)
por: Nikankin, Yaniv, et al.
Publicado: (2024)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
por: Orgad, Hadas, et al.
Publicado: (2024)
por: Orgad, Hadas, et al.
Publicado: (2024)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
por: Feucht, Sheridan, et al.
Publicado: (2026)
por: Feucht, Sheridan, et al.
Publicado: (2026)
Profiling German Text Simplification with Interpretable Model-Fingerprints
por: Klöser, Lars, et al.
Publicado: (2026)
por: Klöser, Lars, et al.
Publicado: (2026)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
por: Sun, Jingyi, et al.
Publicado: (2024)
por: Sun, Jingyi, et al.
Publicado: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
por: Ashuach, Tomer, et al.
Publicado: (2025)
por: Ashuach, Tomer, et al.
Publicado: (2025)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
Evaluating Pixel Language Models on Non-Standardized Languages
por: Muñoz-Ortiz, Alberto, et al.
Publicado: (2024)
por: Muñoz-Ortiz, Alberto, et al.
Publicado: (2024)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
por: Tu, Songjun, et al.
Publicado: (2026)
por: Tu, Songjun, et al.
Publicado: (2026)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
por: Gómez-Rodríguez, Carlos, et al.
Publicado: (2024)
por: Gómez-Rodríguez, Carlos, et al.
Publicado: (2024)
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
por: Bayram, M. Ali, et al.
Publicado: (2024)
por: Bayram, M. Ali, et al.
Publicado: (2024)
Optimizing What We Trust: Reliability-Guided QUBO Selection of Multi-Agent Weak Framing Signals for Arabic Sentiment Prediction
por: Alkhalifa, Rabab
Publicado: (2026)
por: Alkhalifa, Rabab
Publicado: (2026)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
por: Pan, Leyi, et al.
Publicado: (2025)
por: Pan, Leyi, et al.
Publicado: (2025)
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
por: Weigang, Li, et al.
Publicado: (2025)
por: Weigang, Li, et al.
Publicado: (2025)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
por: Yao, Ben, et al.
Publicado: (2025)
por: Yao, Ben, et al.
Publicado: (2025)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
por: Nikankin, Yaniv, et al.
Publicado: (2025)
por: Nikankin, Yaniv, et al.
Publicado: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
por: Fan, Jingxing, et al.
Publicado: (2025)
por: Fan, Jingxing, et al.
Publicado: (2025)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
por: Wang, Yihao, et al.
Publicado: (2026)
por: Wang, Yihao, et al.
Publicado: (2026)
Emotional Sequential Influence Modeling on False Information
por: Naskar, Debashis, et al.
Publicado: (2024)
por: Naskar, Debashis, et al.
Publicado: (2024)
Mixing Times of Glauber Dynamics on Masked Language Models
por: Sana, Suvadip, et al.
Publicado: (2026)
por: Sana, Suvadip, et al.
Publicado: (2026)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
por: Goldin, Gili, et al.
Publicado: (2024)
por: Goldin, Gili, et al.
Publicado: (2024)
Math Natural Language Inference: this should be easy!
por: de Paiva, Valeria, et al.
Publicado: (2025)
por: de Paiva, Valeria, et al.
Publicado: (2025)
Fast Quiet-STaR: Thinking Without Thought Tokens
por: Huang, Wei, et al.
Publicado: (2025)
por: Huang, Wei, et al.
Publicado: (2025)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
por: Wang, Zhilin, et al.
Publicado: (2026)
por: Wang, Zhilin, et al.
Publicado: (2026)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
por: Pan, Leyi, et al.
Publicado: (2025)
por: Pan, Leyi, et al.
Publicado: (2025)
Towards Effective and Efficient Continual Pre-training of Large Language Models
por: Chen, Jie, et al.
Publicado: (2024)
por: Chen, Jie, et al.
Publicado: (2024)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
por: Liu, Aiwei, et al.
Publicado: (2023)
por: Liu, Aiwei, et al.
Publicado: (2023)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
por: Liu, Aiwei, et al.
Publicado: (2024)
por: Liu, Aiwei, et al.
Publicado: (2024)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
por: Kim, Jaemin, et al.
Publicado: (2024)
por: Kim, Jaemin, et al.
Publicado: (2024)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
por: Zhang, Zhehao, et al.
Publicado: (2025)
por: Zhang, Zhehao, et al.
Publicado: (2025)
Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing
por: Fernández-González, Daniel, et al.
Publicado: (2026)
por: Fernández-González, Daniel, et al.
Publicado: (2026)
Parametric Social Identity Injection and Diversification in Public Opinion Simulation
por: Wang, Hexi, et al.
Publicado: (2026)
por: Wang, Hexi, et al.
Publicado: (2026)
Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal
por: Oehri, Markus, et al.
Publicado: (2025)
por: Oehri, Markus, et al.
Publicado: (2025)
Beyond Cosine Similarity
por: Ai, Xinbo
Publicado: (2026)
por: Ai, Xinbo
Publicado: (2026)
AI-assisted German Employment Contract Review: A Benchmark Dataset
por: Wardas, Oliver, et al.
Publicado: (2025)
por: Wardas, Oliver, et al.
Publicado: (2025)
The Superalignment of Superhuman Intelligence with Large Language Models
por: Huang, Minlie, et al.
Publicado: (2024)
por: Huang, Minlie, et al.
Publicado: (2024)
ScoreRAG: A Retrieval-Augmented Generation Framework with Consistency-Relevance Scoring and Structured Summarization for News Generation
por: Lin, Pei-Yun, et al.
Publicado: (2025)
por: Lin, Pei-Yun, et al.
Publicado: (2025)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
por: Park, Seungcheol, et al.
Publicado: (2025)
por: Park, Seungcheol, et al.
Publicado: (2025)
Ejemplares similares
-
Position-aware Automatic Circuit Discovery
por: Haklay, Tal, et al.
Publicado: (2025) -
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
por: Arad, Dana, et al.
Publicado: (2023) -
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
por: Nikankin, Yaniv, et al.
Publicado: (2024) -
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
por: Orgad, Hadas, et al.
Publicado: (2024) -
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
por: Feucht, Sheridan, et al.
Publicado: (2026)