Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chen, Qinhao, He, Linyang, Mesgarani, Nima |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
From Chains to DAGs: Probing the Graph Structure of Reasoning in LLMs
par: Zhong, Tianjun, et autres
Publié: (2026)
par: Zhong, Tianjun, et autres
Publié: (2026)
Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations
par: He, Linyang, et autres
Publié: (2025)
par: He, Linyang, et autres
Publié: (2025)
Transcoders Find Interpretable LLM Feature Circuits
par: Dunefsky, Jacob, et autres
Publié: (2024)
par: Dunefsky, Jacob, et autres
Publié: (2024)
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
par: Wang, Qiaolin, et autres
Publié: (2025)
par: Wang, Qiaolin, et autres
Publié: (2025)
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
par: Draye, Florent, et autres
Publié: (2026)
par: Draye, Florent, et autres
Publié: (2026)
Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence
par: He, Linyang, et autres
Publié: (2024)
par: He, Linyang, et autres
Publié: (2024)
Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement
par: He, Linyang, et autres
Publié: (2025)
par: He, Linyang, et autres
Publié: (2025)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
par: Harrasse, Abir, et autres
Publié: (2025)
par: Harrasse, Abir, et autres
Publié: (2025)
LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches
par: He, Linyang, et autres
Publié: (2026)
par: He, Linyang, et autres
Publié: (2026)
Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
par: Jiang, Xilin, et autres
Publié: (2025)
par: Jiang, Xilin, et autres
Publié: (2025)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
par: Gu, Hao, et autres
Publié: (2025)
par: Gu, Hao, et autres
Publié: (2025)
Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
par: Hatefi, Sayed Mohammad Vakilzadeh, et autres
Publié: (2025)
par: Hatefi, Sayed Mohammad Vakilzadeh, et autres
Publié: (2025)
Contextual Feature Extraction Hierarchies Converge in Large Language Models and the Brain
par: Mischler, Gavin, et autres
Publié: (2024)
par: Mischler, Gavin, et autres
Publié: (2024)
XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
par: He, Linyang, et autres
Publié: (2025)
par: He, Linyang, et autres
Publié: (2025)
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
par: Garg, Ankur, et autres
Publié: (2025)
par: Garg, Ankur, et autres
Publié: (2025)
A cross-species neural foundation model for end-to-end speech decoding
par: Zhang, Yizi, et autres
Publié: (2025)
par: Zhang, Yizi, et autres
Publié: (2025)
Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation
par: Chen, Xinrui, et autres
Publié: (2025)
par: Chen, Xinrui, et autres
Publié: (2025)
Neuro2Semantic: A Transfer Learning Framework for Semantic Reconstruction of Continuous Language from Human Intracranial EEG
par: Shams, Siavash, et autres
Publié: (2025)
par: Shams, Siavash, et autres
Publié: (2025)
Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience
par: Jiang, Xilin, et autres
Publié: (2024)
par: Jiang, Xilin, et autres
Publié: (2024)
Finding Transformer Circuits with Edge Pruning
par: Bhaskar, Adithya, et autres
Publié: (2024)
par: Bhaskar, Adithya, et autres
Publié: (2024)
Dual Perspectives in Emotion Attribution: A Generator-Interpreter Framework for Cross-Cultural Analysis of Emotion in LLMs
par: Turdubaeva, Aizirek, et autres
Publié: (2026)
par: Turdubaeva, Aizirek, et autres
Publié: (2026)
Iterative Layer Pruning for Efficient Translation Inference
par: Moslem, Yasmin, et autres
Publié: (2025)
par: Moslem, Yasmin, et autres
Publié: (2025)
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
par: Kim, Su-Hyeon, et autres
Publié: (2026)
par: Kim, Su-Hyeon, et autres
Publié: (2026)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
par: Anagnostidis, Sotiris, et autres
Publié: (2023)
par: Anagnostidis, Sotiris, et autres
Publié: (2023)
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
par: Li, Yinghao Aaron, et autres
Publié: (2024)
par: Li, Yinghao Aaron, et autres
Publié: (2024)
A Simple Linear Patch Revives Layer-Pruned Large Language Models
par: Chen, Xinrui, et autres
Publié: (2025)
par: Chen, Xinrui, et autres
Publié: (2025)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
par: Chatzoudis, Gerasimos, et autres
Publié: (2026)
par: Chatzoudis, Gerasimos, et autres
Publié: (2026)
Automated Interpretability and Feature Discovery in Language Models with Agents
par: Marin-Llobet, Arnau, et autres
Publié: (2026)
par: Marin-Llobet, Arnau, et autres
Publié: (2026)
Towards Building Efficient Sentence BERT Models using Layer Pruning
par: Shelke, Anushka, et autres
Publié: (2024)
par: Shelke, Anushka, et autres
Publié: (2024)
CFSP: An Efficient Structured Pruning Framework for LLMs with Coarse-to-Fine Activation Information
par: Wang, Yuxin, et autres
Publié: (2024)
par: Wang, Yuxin, et autres
Publié: (2024)
NutePrune: Efficient Progressive Pruning with Numerous Teachers for Large Language Models
par: Li, Shengrui, et autres
Publié: (2024)
par: Li, Shengrui, et autres
Publié: (2024)
All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs
par: Chen, Xi, et autres
Publié: (2026)
par: Chen, Xi, et autres
Publié: (2026)
E$^3$-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
par: Yuan, Tao, et autres
Publié: (2025)
par: Yuan, Tao, et autres
Publié: (2025)
Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
par: Wu, Junkai, et autres
Publié: (2024)
par: Wu, Junkai, et autres
Publié: (2024)
Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation
par: Jantsch, Lasse Marten, et autres
Publié: (2026)
par: Jantsch, Lasse Marten, et autres
Publié: (2026)
Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
par: Liu, Deyuan, et autres
Publié: (2024)
par: Liu, Deyuan, et autres
Publié: (2024)
Swift Cross-Dataset Pruning: Enhancing Fine-Tuning Efficiency in Natural Language Understanding
par: Nguyen, Binh-Nguyen, et autres
Publié: (2025)
par: Nguyen, Binh-Nguyen, et autres
Publié: (2025)
Contextual Compression Encoding for Large Language Models: A Novel Framework for Multi-Layered Parameter Space Pruning
par: Schmitt, Barnaby, et autres
Publié: (2025)
par: Schmitt, Barnaby, et autres
Publié: (2025)
EHRSummarizer: A Privacy-Aware, FHIR-Native Reference Architecture for Source-Grounded EHR Summarization
par: Kazemzadeh, Houman, et autres
Publié: (2026)
par: Kazemzadeh, Houman, et autres
Publié: (2026)
Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models
par: Damianos, Dimitrios, et autres
Publié: (2026)
par: Damianos, Dimitrios, et autres
Publié: (2026)
Documents similaires
-
From Chains to DAGs: Probing the Graph Structure of Reasoning in LLMs
par: Zhong, Tianjun, et autres
Publié: (2026) -
Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations
par: He, Linyang, et autres
Publié: (2025) -
Transcoders Find Interpretable LLM Feature Circuits
par: Dunefsky, Jacob, et autres
Publié: (2024) -
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
par: Wang, Qiaolin, et autres
Publié: (2025) -
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
par: Draye, Florent, et autres
Publié: (2026)