Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation
Fuente:
arXiv
Salvato in:
| Autori principali: | Jantsch, Lasse Marten, Koh, Dong-Jae, Lee, Seonghyeon, Suh, Young-Kyoon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Depth Registers Unlock W4A4 on SwiGLU: A Reader/Generator Decomposition
di: Liu, Ziyang
Pubblicazione: (2026)
di: Liu, Ziyang
Pubblicazione: (2026)
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
LoLA-SpecViT: Local Attention SwiGLU Vision Transformer with LoRA for Hyperspectral Imaging
di: Zidi, Fadi Abdeladhim, et al.
Pubblicazione: (2025)
di: Zidi, Fadi Abdeladhim, et al.
Pubblicazione: (2025)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
di: Lau, Tim Tsz-Kit, et al.
Pubblicazione: (2026)
di: Lau, Tim Tsz-Kit, et al.
Pubblicazione: (2026)
Hidden Heroes and Gradient Bloats: Layer-Wise Redundancy Inverts Attribution in Transformers
di: Ye, Donald
Pubblicazione: (2026)
di: Ye, Donald
Pubblicazione: (2026)
GLU Attention Improve Transformer
di: Wang, Zehao
Pubblicazione: (2025)
di: Wang, Zehao
Pubblicazione: (2025)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
di: Achtibat, Reduan, et al.
Pubblicazione: (2024)
di: Achtibat, Reduan, et al.
Pubblicazione: (2024)
Table Transformers for Imputing Textual Attributes
di: Wei, Ting-Ruen, et al.
Pubblicazione: (2024)
di: Wei, Ting-Ruen, et al.
Pubblicazione: (2024)
Explanation Regularisation through the Lens of Attributions
di: Ferreira, Pedro, et al.
Pubblicazione: (2024)
di: Ferreira, Pedro, et al.
Pubblicazione: (2024)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
di: Gerstner, Sebastian, et al.
Pubblicazione: (2026)
di: Gerstner, Sebastian, et al.
Pubblicazione: (2026)
Dual Perspectives in Emotion Attribution: A Generator-Interpreter Framework for Cross-Cultural Analysis of Emotion in LLMs
di: Turdubaeva, Aizirek, et al.
Pubblicazione: (2026)
di: Turdubaeva, Aizirek, et al.
Pubblicazione: (2026)
SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
Approximate Attributions for Off-the-Shelf Siamese Transformers
di: Möller, Lucas, et al.
Pubblicazione: (2024)
di: Möller, Lucas, et al.
Pubblicazione: (2024)
Reliable, Adaptable, and Attributable Language Models with Retrieval
di: Asai, Akari, et al.
Pubblicazione: (2024)
di: Asai, Akari, et al.
Pubblicazione: (2024)
Attribution functionalism
di: Mark Phelan
Pubblicazione: (2025)
di: Mark Phelan
Pubblicazione: (2025)
Efficient Text-Attributed Graph Learning through Selective Annotation and Graph Alignment
di: Xie, Huanyi, et al.
Pubblicazione: (2025)
di: Xie, Huanyi, et al.
Pubblicazione: (2025)
Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution
di: Chen, Qinhao, et al.
Pubblicazione: (2026)
di: Chen, Qinhao, et al.
Pubblicazione: (2026)
LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
di: Bansal, Harsh Vardhan
Pubblicazione: (2025)
di: Bansal, Harsh Vardhan
Pubblicazione: (2025)
Advancing Attribution-Based Neural Network Explainability through Relative Absolute Magnitude Layer-Wise Relevance Propagation and Multi-Component Evaluation
di: Vukadin, Davor, et al.
Pubblicazione: (2024)
di: Vukadin, Davor, et al.
Pubblicazione: (2024)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
Multi-Attribute Steering of Language Models via Targeted Intervention
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
VISTA: Visualization of Token Attribution via Efficient Analysis
di: Ahmed, Syed, et al.
Pubblicazione: (2026)
di: Ahmed, Syed, et al.
Pubblicazione: (2026)
AttributionBench: How Hard is Automatic Attribution Evaluation?
di: Li, Yifei, et al.
Pubblicazione: (2024)
di: Li, Yifei, et al.
Pubblicazione: (2024)
SwiLTra-Bench: The Swiss Legal Translation Benchmark
di: Niklaus, Joel, et al.
Pubblicazione: (2025)
di: Niklaus, Joel, et al.
Pubblicazione: (2025)
One Arrow, Many Targets: Probing LLMs for Multi-Attribute Controllable Text Summarization
di: Roy, Tathagato, et al.
Pubblicazione: (2024)
di: Roy, Tathagato, et al.
Pubblicazione: (2024)
Bias Analysis and Mitigation through Protected Attribute Detection and Regard Classification
di: Udagawa, Takuma, et al.
Pubblicazione: (2025)
di: Udagawa, Takuma, et al.
Pubblicazione: (2025)
Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
di: Proietti, Michela, et al.
Pubblicazione: (2025)
di: Proietti, Michela, et al.
Pubblicazione: (2025)
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
di: Draye, Florent, et al.
Pubblicazione: (2026)
di: Draye, Florent, et al.
Pubblicazione: (2026)
Advancing Large Language Model Attribution through Self-Improving
di: Huang, Lei, et al.
Pubblicazione: (2024)
di: Huang, Lei, et al.
Pubblicazione: (2024)
Learning to Attribute with Attention
di: Cohen-Wang, Benjamin, et al.
Pubblicazione: (2025)
di: Cohen-Wang, Benjamin, et al.
Pubblicazione: (2025)
Evaluating the Smooth Control of Attribute Intensity in Text Generation with LLMs
di: Zhou, Shang, et al.
Pubblicazione: (2024)
di: Zhou, Shang, et al.
Pubblicazione: (2024)
Efficient Estimation of Kernel Surrogate Models for Task Attribution
di: Zhang, Zhenshuo, et al.
Pubblicazione: (2026)
di: Zhang, Zhenshuo, et al.
Pubblicazione: (2026)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
di: Mihaila, George
Pubblicazione: (2026)
di: Mihaila, George
Pubblicazione: (2026)
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
di: Song, Maojia, et al.
Pubblicazione: (2024)
di: Song, Maojia, et al.
Pubblicazione: (2024)
MentalAgora: A Gateway to Advanced Personalized Care in Mental Health through Multi-Agent Debating and Attribute Control
di: Lee, Yeonji, et al.
Pubblicazione: (2024)
di: Lee, Yeonji, et al.
Pubblicazione: (2024)
KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark
di: Jang, Seongbo, et al.
Pubblicazione: (2024)
di: Jang, Seongbo, et al.
Pubblicazione: (2024)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
Improving Attributed Long-form Question Answering with Intent Awareness
di: Zhao, Xinran, et al.
Pubblicazione: (2026)
di: Zhao, Xinran, et al.
Pubblicazione: (2026)
Quantifying Misattribution Unfairness in Authorship Attribution
di: Alipoormolabashi, Pegah, et al.
Pubblicazione: (2025)
di: Alipoormolabashi, Pegah, et al.
Pubblicazione: (2025)
On the Feasibility of In-Context Probing for Data Attribution
di: Jiao, Cathy, et al.
Pubblicazione: (2024)
di: Jiao, Cathy, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Depth Registers Unlock W4A4 on SwiGLU: A Reader/Generator Decomposition
di: Liu, Ziyang
Pubblicazione: (2026) -
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
di: Yadav, Sarthak, et al.
Pubblicazione: (2025) -
LoLA-SpecViT: Local Attention SwiGLU Vision Transformer with LoRA for Hyperspectral Imaging
di: Zidi, Fadi Abdeladhim, et al.
Pubblicazione: (2025) -
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
di: Lau, Tim Tsz-Kit, et al.
Pubblicazione: (2026) -
Hidden Heroes and Gradient Bloats: Layer-Wise Redundancy Inverts Attribution in Transformers
di: Ye, Donald
Pubblicazione: (2026)