Hidden Heroes and Gradient Bloats: Layer-Wise Redundancy Inverts Attribution in Transformers
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Ye, Donald |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer Models
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2026)
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2026)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026)
LPCD: Unified Framework from Layer-Wise to Submodule Quantization
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025)
Layer by Layer: Uncovering Hidden Representations in Language Models
von: Skean, Oscar, et al.
Veröffentlicht: (2025)
von: Skean, Oscar, et al.
Veröffentlicht: (2025)
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering
von: Zhou, Yefan, et al.
Veröffentlicht: (2026)
von: Zhou, Yefan, et al.
Veröffentlicht: (2026)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
von: Achtibat, Reduan, et al.
Veröffentlicht: (2024)
von: Achtibat, Reduan, et al.
Veröffentlicht: (2024)
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
von: Sim, Woo Seob, et al.
Veröffentlicht: (2026)
von: Sim, Woo Seob, et al.
Veröffentlicht: (2026)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
von: Kim, Minkyu, et al.
Veröffentlicht: (2026)
No Free Swap: Protocol-Dependent Layer Redundancy in Transformers
von: Garcia, Gabriel
Veröffentlicht: (2026)
von: Garcia, Gabriel
Veröffentlicht: (2026)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
von: Li, Xing, et al.
Veröffentlicht: (2025)
von: Li, Xing, et al.
Veröffentlicht: (2025)
Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity
von: Jiang, Jiachen, et al.
Veröffentlicht: (2024)
von: Jiang, Jiachen, et al.
Veröffentlicht: (2024)
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
Can LLMs Convert Graphs to Text-Attributed Graphs?
von: Wang, Zehong, et al.
Veröffentlicht: (2024)
von: Wang, Zehong, et al.
Veröffentlicht: (2024)
Dual Path Attribution: Efficient Attribution for SwiGLU-Transformers through Layer-Wise Target Propagation
von: Jantsch, Lasse Marten, et al.
Veröffentlicht: (2026)
von: Jantsch, Lasse Marten, et al.
Veröffentlicht: (2026)
Think Clearly: Improving Reasoning via Redundant Token Pruning
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
von: Kim, Jeonghoon, et al.
Veröffentlicht: (2025)
von: Kim, Jeonghoon, et al.
Veröffentlicht: (2025)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
von: Guo, Song, et al.
Veröffentlicht: (2024)
von: Guo, Song, et al.
Veröffentlicht: (2024)
GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
von: Ye, Tian, et al.
Veröffentlicht: (2024)
von: Ye, Tian, et al.
Veröffentlicht: (2024)
AttributionBench: How Hard is Automatic Attribution Evaluation?
von: Li, Yifei, et al.
Veröffentlicht: (2024)
von: Li, Yifei, et al.
Veröffentlicht: (2024)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
von: Kamigaito, Hidetaka, et al.
Veröffentlicht: (2025)
von: Kamigaito, Hidetaka, et al.
Veröffentlicht: (2025)
What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
von: He, Junhui, et al.
Veröffentlicht: (2024)
von: He, Junhui, et al.
Veröffentlicht: (2024)
SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
von: Czakó, Patrik, et al.
Veröffentlicht: (2025)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
von: Kapadia, Shashank, et al.
Veröffentlicht: (2026)
von: Kapadia, Shashank, et al.
Veröffentlicht: (2026)
Efficient Reasoning with Hidden Thinking
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
von: Sung, Yi-Lin, et al.
Veröffentlicht: (2023)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
von: Lu, Liming, et al.
Veröffentlicht: (2026)
von: Lu, Liming, et al.
Veröffentlicht: (2026)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
von: Fu, Zihao, et al.
Veröffentlicht: (2025)
Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Rongzhi, et al.
Veröffentlicht: (2025)
GoRA: Gradient-driven Adaptive Low Rank Adaptation
von: He, Haonan, et al.
Veröffentlicht: (2025)
von: He, Haonan, et al.
Veröffentlicht: (2025)
LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models
von: Li, Guangyan, et al.
Veröffentlicht: (2024)
von: Li, Guangyan, et al.
Veröffentlicht: (2024)
Model Assembly Learning with Heterogeneous Layer Weight Merging
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2025)
$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction
von: Du, Zhenbang, et al.
Veröffentlicht: (2026)
von: Du, Zhenbang, et al.
Veröffentlicht: (2026)
QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance
von: Saha, Binita, et al.
Veröffentlicht: (2025)
von: Saha, Binita, et al.
Veröffentlicht: (2025)
FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
von: Zou, Heming, et al.
Veröffentlicht: (2025)
von: Zou, Heming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Explainable AI: Context-Aware Layer-Wise Integrated Gradients for Explaining Transformer Models
von: Mersha, Melkamu Abay, et al.
Veröffentlicht: (2026) -
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
von: Taniguchi, Rei, et al.
Veröffentlicht: (2026) -
LPCD: Unified Framework from Layer-Wise to Submodule Quantization
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2025) -
Layer by Layer: Uncovering Hidden Representations in Language Models
von: Skean, Oscar, et al.
Veröffentlicht: (2025) -
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)