Saved in:
| Main Authors: | He, Zhengfu, Wang, Junxuan, Lin, Rui, Ge, Xuyang, Shu, Wentao, Tang, Qiong, Zhang, Junping, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.20938 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025)
by: Wang, Junxuan, et al.
Published: (2025)
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
by: Ge, Xuyang, et al.
Published: (2024)
by: Ge, Xuyang, et al.
Published: (2024)
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
by: Wang, Junxuan, et al.
Published: (2024)
by: Wang, Junxuan, et al.
Published: (2024)
Tracing the Thought of a Grandmaster-level Chess-Playing Transformer
by: Lin, Rui, et al.
Published: (2026)
by: Lin, Rui, et al.
Published: (2026)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
by: He, Zhengfu, et al.
Published: (2024)
by: He, Zhengfu, et al.
Published: (2024)
Evolution of Concepts in Language Model Pre-Training
by: Ge, Xuyang, et al.
Published: (2025)
by: Ge, Xuyang, et al.
Published: (2025)
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
by: He, Zhengfu, et al.
Published: (2024)
by: He, Zhengfu, et al.
Published: (2024)
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
by: Zhou, Guancheng, et al.
Published: (2026)
by: Zhou, Guancheng, et al.
Published: (2026)
LoLA: Low-Rank Linear Attention With Sparse Caching
by: McDermott, Luke, et al.
Published: (2025)
by: McDermott, Luke, et al.
Published: (2025)
Sparse Attention Decomposition Applied to Circuit Tracing
by: Franco, Gabriel, et al.
Published: (2024)
by: Franco, Gabriel, et al.
Published: (2024)
MAP: Revisiting Weight Decomposition for Low-Rank Adaptation
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
by: Lv, Kai, et al.
Published: (2023)
by: Lv, Kai, et al.
Published: (2023)
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
by: Cho, Yoonjun, et al.
Published: (2025)
by: Cho, Yoonjun, et al.
Published: (2025)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
DropLoRA: Sparse Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
by: Zhang, Haojie
Published: (2025)
by: Zhang, Haojie
Published: (2025)
Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO
by: Zeng, Zhiyuan, et al.
Published: (2026)
by: Zeng, Zhiyuan, et al.
Published: (2026)
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
by: Li, Yun, et al.
Published: (2023)
by: Li, Yun, et al.
Published: (2023)
Efficient Low Rank Attention for Long-Context Inference in Large Language Models
by: Li, Tenghui, et al.
Published: (2025)
by: Li, Tenghui, et al.
Published: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
by: Tang, Chuanyu, et al.
Published: (2024)
by: Tang, Chuanyu, et al.
Published: (2024)
A3 : an Analytical Low-Rank Approximation Framework for Attention
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action Models
by: Liu, Peiju, et al.
Published: (2026)
by: Liu, Peiju, et al.
Published: (2026)
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
by: Zhang, Mozhi, et al.
Published: (2025)
by: Zhang, Mozhi, et al.
Published: (2025)
Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
by: Li, Shiwei, et al.
Published: (2025)
by: Li, Shiwei, et al.
Published: (2025)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
by: He, Mutian, et al.
Published: (2025)
by: He, Mutian, et al.
Published: (2025)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
by: Leng, Jiaqi, et al.
Published: (2025)
by: Leng, Jiaqi, et al.
Published: (2025)
Towards Understanding the Robustness of Sparse Autoencoders
by: Saiyed, Ahson, et al.
Published: (2026)
by: Saiyed, Ahson, et al.
Published: (2026)
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
by: Erden, Caner
Published: (2025)
by: Erden, Caner
Published: (2025)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023)
by: Gupta, Ahan, et al.
Published: (2023)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
Prism: Spectral-Aware Block-Sparse Attention
by: Wang, Xinghao, et al.
Published: (2026)
by: Wang, Xinghao, et al.
Published: (2026)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
Revisiting Graph-Tokenizing Large Language Models: A Systematic Evaluation of Graph Token Understanding
by: Zhang, Zhongjian, et al.
Published: (2026)
by: Zhang, Zhongjian, et al.
Published: (2026)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
by: Zhang, Wenzheng, et al.
Published: (2026)
by: Zhang, Wenzheng, et al.
Published: (2026)
EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition
by: Nasiri, Hamid, et al.
Published: (2025)
by: Nasiri, Hamid, et al.
Published: (2025)
Similar Items
-
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025) -
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
by: Ge, Xuyang, et al.
Published: (2024) -
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
by: Wang, Junxuan, et al.
Published: (2024) -
Tracing the Thought of a Grandmaster-level Chess-Playing Transformer
by: Lin, Rui, et al.
Published: (2026) -
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
by: He, Zhengfu, et al.
Published: (2024)