HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Jiuding, Huang, Jing, Baskaran, Sidharth, D'Oosterlinck, Karel, Potts, Christopher, Sklar, Michael, Geiger, Atticus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HyperSteer: Activation Steering at Scale with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
Updating CLIP to Prefer Descriptions Over Captions
by: Zur, Amir, et al.
Published: (2024)
by: Zur, Amir, et al.
Published: (2024)
In-Context Learning for Extreme Multi-Label Classification
by: D'Oosterlinck, Karel, et al.
Published: (2024)
by: D'Oosterlinck, Karel, et al.
Published: (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
by: Huang, Jing, et al.
Published: (2024)
by: Huang, Jing, et al.
Published: (2024)
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
by: Wu, Zhengxuan, et al.
Published: (2023)
by: Wu, Zhengxuan, et al.
Published: (2023)
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
by: D'Oosterlinck, Karel, et al.
Published: (2024)
by: D'Oosterlinck, Karel, et al.
Published: (2024)
A Reply to Makelov et al. (2023)'s "Interpretability Illusion" Arguments
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
Constructing Interpretable Features from Compositional Neuron Groups
by: Shafran, Or, et al.
Published: (2025)
by: Shafran, Or, et al.
Published: (2025)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
by: She, Jingyuan Selena, et al.
Published: (2023)
by: She, Jingyuan Selena, et al.
Published: (2023)
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
Composing Policy Gradients and Prompt Optimization for Language Model Programs
by: Ziems, Noah, et al.
Published: (2025)
by: Ziems, Noah, et al.
Published: (2025)
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
ReFT: Representation Finetuning for Language Models
by: Wu, Zhengxuan, et al.
Published: (2024)
by: Wu, Zhengxuan, et al.
Published: (2024)
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
Hyper-CL: Conditioning Sentence Representations with Hypernetworks
by: Yoo, Young Hyun, et al.
Published: (2024)
by: Yoo, Young Hyun, et al.
Published: (2024)
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
MIB: A Mechanistic Interpretability Benchmark
by: Mueller, Aaron, et al.
Published: (2025)
by: Mueller, Aaron, et al.
Published: (2025)
Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations
by: Geiger, Atticus, et al.
Published: (2023)
by: Geiger, Atticus, et al.
Published: (2023)
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
How Causal Abstraction Underpins Computational Explanation
by: Geiger, Atticus, et al.
Published: (2025)
by: Geiger, Atticus, et al.
Published: (2025)
How Do Transformers Learn Variable Binding in Symbolic Programs?
by: Wu, Yiwei, et al.
Published: (2025)
by: Wu, Yiwei, et al.
Published: (2025)
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
FLRT: Fluent Student-Teacher Redteaming
by: Thompson, T. Ben, et al.
Published: (2024)
by: Thompson, T. Ben, et al.
Published: (2024)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
by: Zur, Amir, et al.
Published: (2025)
by: Zur, Amir, et al.
Published: (2025)
HyperEdit: Unlocking Instruction-based Text Editing in LLMs via Hypernetworks
by: Zeng, Yiming, et al.
Published: (2025)
by: Zeng, Yiming, et al.
Published: (2025)
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
by: Wu, Zhengxuan, et al.
Published: (2023)
by: Wu, Zhengxuan, et al.
Published: (2023)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
by: Li, Yingting, et al.
Published: (2024)
by: Li, Yingting, et al.
Published: (2024)
Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction
by: Puyin, Li, et al.
Published: (2026)
by: Puyin, Li, et al.
Published: (2026)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
by: Wang, Atticus, et al.
Published: (2025)
by: Wang, Atticus, et al.
Published: (2025)
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry
by: Shafran, Or, et al.
Published: (2026)
by: Shafran, Or, et al.
Published: (2026)
Activation Steering via Generative Causal Mediation
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
by: Sankaranarayanan, Aruna, et al.
Published: (2026)
HyperLoader: Integrating Hypernetwork-Based LoRA and Adapter Layers into Multi-Task Transformers for Sequence Labelling
by: Ortiz-Barajas, Jesus-German, et al.
Published: (2024)
by: Ortiz-Barajas, Jesus-German, et al.
Published: (2024)
Fluent dreaming for language models
by: Thompson, T. Ben, et al.
Published: (2024)
by: Thompson, T. Ben, et al.
Published: (2024)
Demystifying Verbatim Memorization in Large Language Models
by: Huang, Jing, et al.
Published: (2024)
by: Huang, Jing, et al.
Published: (2024)
HyperAdaLoRA: Accelerating LoRA Rank Allocation During Training via Hypernetworks without Sacrificing Performance
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research
by: Trott, Sean
Published: (2025)
by: Trott, Sean
Published: (2025)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
by: Sun, Alan, et al.
Published: (2026)
by: Sun, Alan, et al.
Published: (2026)
VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback
by: Zhang, Guoxi, et al.
Published: (2024)
by: Zhang, Guoxi, et al.
Published: (2024)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
Similar Items
-
HyperSteer: Activation Steering at Scale with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025) -
Updating CLIP to Prefer Descriptions Over Captions
by: Zur, Amir, et al.
Published: (2024) -
In-Context Learning for Extreme Multi-Label Classification
by: D'Oosterlinck, Karel, et al.
Published: (2024) -
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
by: Huang, Jing, et al.
Published: (2024) -
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
by: Wu, Zhengxuan, et al.
Published: (2023)