Because we have LLMs, we Can and Should Pursue Agentic Interpretability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Been, Hewitt, John, Nanda, Neel, Fiedel, Noah, Tafjord, Oyvind |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
We Can't Understand AI Using our Existing Vocabulary
von: Hewitt, John, et al.
Veröffentlicht: (2025)
von: Hewitt, John, et al.
Veröffentlicht: (2025)
Neologism Learning for Controllability and Self-Verbalization
von: Hewitt, John, et al.
Veröffentlicht: (2025)
von: Hewitt, John, et al.
Veröffentlicht: (2025)
Digital Socrates: Evaluating LLMs through Explanation Critiques
von: Gu, Yuling, et al.
Veröffentlicht: (2023)
von: Gu, Yuling, et al.
Veröffentlicht: (2023)
BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability
von: Clark, Peter, et al.
Veröffentlicht: (2023)
von: Clark, Peter, et al.
Veröffentlicht: (2023)
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
von: Smit, Andries, et al.
Veröffentlicht: (2023)
von: Smit, Andries, et al.
Veröffentlicht: (2023)
Can we forget how we learned? Doxastic redundancy in iterated belief revision
von: Liberatore, Paolo
Veröffentlicht: (2024)
von: Liberatore, Paolo
Veröffentlicht: (2024)
Features have life history. And we should care
von: Stecher, Philipp, et al.
Veröffentlicht: (2026)
von: Stecher, Philipp, et al.
Veröffentlicht: (2026)
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
von: Liu, Xinyu, et al.
Veröffentlicht: (2024)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
Can we Evaluate RAGs with Synthetic Data?
von: van Elburg, Jonas, et al.
Veröffentlicht: (2025)
von: van Elburg, Jonas, et al.
Veröffentlicht: (2025)
When AI reviews science: Can we trust the referee?
von: Wang, Jialiang, et al.
Veröffentlicht: (2026)
von: Wang, Jialiang, et al.
Veröffentlicht: (2026)
Can we automatize scientific discovery in the cognitive sciences?
von: Jagadish, Akshay K., et al.
Veröffentlicht: (2026)
von: Jagadish, Akshay K., et al.
Veröffentlicht: (2026)
Can we trust the evaluation on ChatGPT?
von: Aiyappa, Rachith, et al.
Veröffentlicht: (2023)
von: Aiyappa, Rachith, et al.
Veröffentlicht: (2023)
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
Explorations of Self-Repair in Language Models
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
OLMES: A Standard for Language Model Evaluations
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
von: Gu, Yuling, et al.
Veröffentlicht: (2024)
How Well Do Models Follow Their Constitutions?
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
LitLLMs, LLMs for Literature Review: Are we there yet?
von: Agarwal, Shubham, et al.
Veröffentlicht: (2024)
von: Agarwal, Shubham, et al.
Veröffentlicht: (2024)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
von: Minder, Julian, et al.
Veröffentlicht: (2025)
von: Minder, Julian, et al.
Veröffentlicht: (2025)
Can we only use guideline instead of shot in prompt?
von: Chen, Jiaxiang, et al.
Veröffentlicht: (2024)
von: Chen, Jiaxiang, et al.
Veröffentlicht: (2024)
BatchTopK Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
von: Bussmann, Bart, et al.
Veröffentlicht: (2024)
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
von: Tzachristas, Ioannis, et al.
Veröffentlicht: (2025)
von: Tzachristas, Ioannis, et al.
Veröffentlicht: (2025)
Can we ease the Injectivity Bottleneck on Lorentzian Manifolds for Graph Neural Networks?
von: Srinivasan, Srinitish, et al.
Veröffentlicht: (2025)
von: Srinivasan, Srinitish, et al.
Veröffentlicht: (2025)
Because we're here
von: Susskind, Leonard
Veröffentlicht: (2005)
von: Susskind, Leonard
Veröffentlicht: (2005)
Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
von: Albers, Nele, et al.
Veröffentlicht: (2025)
von: Albers, Nele, et al.
Veröffentlicht: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
von: Maar, Jim, et al.
Veröffentlicht: (2026)
von: Maar, Jim, et al.
Veröffentlicht: (2026)
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
von: Agarwal, Utkarsh, et al.
Veröffentlicht: (2024)
von: Agarwal, Utkarsh, et al.
Veröffentlicht: (2024)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
von: Casademunt, Helena, et al.
Veröffentlicht: (2026)
GPT-ology, Computational Models, Silicon Sampling: How should we think about LLMs in Cognitive Science?
von: Ong, Desmond C.
Veröffentlicht: (2024)
von: Ong, Desmond C.
Veröffentlicht: (2024)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
von: Bussmann, Bart, et al.
Veröffentlicht: (2025)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
Subliminal Learning Is Steering Vector Distillation
von: Blank, Camila, et al.
Veröffentlicht: (2026)
von: Blank, Camila, et al.
Veröffentlicht: (2026)
Why should we ever automate moral decision making?
von: Conitzer, Vincent
Veröffentlicht: (2024)
von: Conitzer, Vincent
Veröffentlicht: (2024)
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
von: Hua, Tim Tian, et al.
Veröffentlicht: (2025)
von: Hua, Tim Tian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
We Can't Understand AI Using our Existing Vocabulary
von: Hewitt, John, et al.
Veröffentlicht: (2025) -
Neologism Learning for Controllability and Self-Verbalization
von: Hewitt, John, et al.
Veröffentlicht: (2025) -
Digital Socrates: Evaluating LLMs through Explanation Critiques
von: Gu, Yuling, et al.
Veröffentlicht: (2023) -
BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability
von: Clark, Peter, et al.
Veröffentlicht: (2023) -
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
von: Smit, Andries, et al.
Veröffentlicht: (2023)