Scaling Attention via Feature Sparsity
Fuente:
arXiv
Guardado en:
| Autores principales: | Xie, Yan, Wen, Tiansheng, Huang, Tangda, Chen, Bo, You, Chenyu, Jegelka, Stefanie, Wang, Yifei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
por: Guo, Lixuan, et al.
Publicado: (2026)
por: Guo, Lixuan, et al.
Publicado: (2026)
CSRv2: Unlocking Ultra-Sparse Embeddings
por: Guo, Lixuan, et al.
Publicado: (2026)
por: Guo, Lixuan, et al.
Publicado: (2026)
Route Experts by Sequence, not by Token
por: Wen, Tiansheng, et al.
Publicado: (2025)
por: Wen, Tiansheng, et al.
Publicado: (2025)
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
por: Wen, Tiansheng, et al.
Publicado: (2025)
por: Wen, Tiansheng, et al.
Publicado: (2025)
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
por: Zhang, Qi, et al.
Publicado: (2024)
por: Zhang, Qi, et al.
Publicado: (2024)
How to Craft Backdoors with Unlabeled Data Alone?
por: Wang, Yifei, et al.
Publicado: (2024)
por: Wang, Yifei, et al.
Publicado: (2024)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
por: Li, Ang, et al.
Publicado: (2025)
por: Li, Ang, et al.
Publicado: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
por: Guo, Xiaojun, et al.
Publicado: (2025)
por: Guo, Xiaojun, et al.
Publicado: (2025)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
por: Hu, Junhao, et al.
Publicado: (2025)
por: Hu, Junhao, et al.
Publicado: (2025)
Geometric and Dynamic Scaling in Deep Transformers
por: Su, Haoran, et al.
Publicado: (2026)
por: Su, Haoran, et al.
Publicado: (2026)
When More is Less: Understanding Chain-of-Thought Length in LLMs
por: Wu, Yuyang, et al.
Publicado: (2025)
por: Wu, Yuyang, et al.
Publicado: (2025)
Contrastive Factor Analysis
por: Duan, Zhibin, et al.
Publicado: (2024)
por: Duan, Zhibin, et al.
Publicado: (2024)
Geometric Algorithms for Neural Combinatorial Optimization with Constraints
por: Karalias, Nikolaos, et al.
Publicado: (2025)
por: Karalias, Nikolaos, et al.
Publicado: (2025)
Learning with Exact Invariances in Polynomial Time
por: Soleymani, Ashkan, et al.
Publicado: (2025)
por: Soleymani, Ashkan, et al.
Publicado: (2025)
Survey on Generalization Theory for Graph Neural Networks
por: Vasileiou, Antonis, et al.
Publicado: (2025)
por: Vasileiou, Antonis, et al.
Publicado: (2025)
An Information Criterion for Controlled Disentanglement of Multimodal Data
por: Wang, Chenyu, et al.
Publicado: (2024)
por: Wang, Chenyu, et al.
Publicado: (2024)
PRISM: Mitigating EHR Data Sparsity via Learning from Missing Feature Calibrated Prototype Patient Representations
por: Zhu, Yinghao, et al.
Publicado: (2023)
por: Zhu, Yinghao, et al.
Publicado: (2023)
HashAttention: Semantic Sparsity for Faster Inference
por: Desai, Aditya, et al.
Publicado: (2024)
por: Desai, Aditya, et al.
Publicado: (2024)
Learning Diffusion Models with Flexible Representation Guidance
por: Wang, Chenyu, et al.
Publicado: (2025)
por: Wang, Chenyu, et al.
Publicado: (2025)
From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
por: Ren, Yuxin, et al.
Publicado: (2026)
por: Ren, Yuxin, et al.
Publicado: (2026)
Multi-View Graph Feature Propagation for Privacy Preservation and Feature Sparsity
por: Harari, Etzion, et al.
Publicado: (2025)
por: Harari, Etzion, et al.
Publicado: (2025)
Fairness Aware Reward Optimization
por: Choi, Ching Lam, et al.
Publicado: (2026)
por: Choi, Ching Lam, et al.
Publicado: (2026)
The Empirical Impact of Neural Parameter Symmetries, or Lack Thereof
por: Lim, Derek, et al.
Publicado: (2024)
por: Lim, Derek, et al.
Publicado: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
por: Gatmiry, Khashayar, et al.
Publicado: (2024)
por: Gatmiry, Khashayar, et al.
Publicado: (2024)
On the Stability of Expressive Positional Encodings for Graphs
por: Huang, Yinan, et al.
Publicado: (2023)
por: Huang, Yinan, et al.
Publicado: (2023)
Understanding the Role of Equivariance in Self-supervised Learning
por: Wang, Yifei, et al.
Publicado: (2024)
por: Wang, Yifei, et al.
Publicado: (2024)
Enhancing Imbalanced Node Classification via Curriculum-Guided Feature Learning and Three-Stage Attention Network
por: Fofanah, Abdul Joseph, et al.
Publicado: (2026)
por: Fofanah, Abdul Joseph, et al.
Publicado: (2026)
A Non-negative VAE:the Generalized Gamma Belief Network
por: Duan, Zhibin, et al.
Publicado: (2024)
por: Duan, Zhibin, et al.
Publicado: (2024)
Learning Linear Attention in Polynomial Time
por: Yau, Morris, et al.
Publicado: (2024)
por: Yau, Morris, et al.
Publicado: (2024)
Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning
por: Ma, Guozheng, et al.
Publicado: (2025)
por: Ma, Guozheng, et al.
Publicado: (2025)
Empty SPACE: Cross-Attention Sparsity for Concept Erasure in Diffusion Models
por: Novello, Nicola, et al.
Publicado: (2026)
por: Novello, Nicola, et al.
Publicado: (2026)
FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
por: Gao, Yifei, et al.
Publicado: (2025)
por: Gao, Yifei, et al.
Publicado: (2025)
ScenGAN: Attention-Intensive Generative Model for Uncertainty-Aware Renewable Scenario Forecasting
por: Wu, Yifei, et al.
Publicado: (2025)
por: Wu, Yifei, et al.
Publicado: (2025)
Post-Training Sparse Attention with Double Sparsity
por: Yang, Shuo, et al.
Publicado: (2024)
por: Yang, Shuo, et al.
Publicado: (2024)
PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity
por: Kim, Kwanyoung, et al.
Publicado: (2025)
por: Kim, Kwanyoung, et al.
Publicado: (2025)
Improving MLLM Training Efficiency via Stage-Aware Sparsity
por: Shi, Kean, et al.
Publicado: (2025)
por: Shi, Kean, et al.
Publicado: (2025)
Improving Decision Sparsity
por: Sun, Yiyang, et al.
Publicado: (2024)
por: Sun, Yiyang, et al.
Publicado: (2024)
MVCL-DAF++: Enhancing Multimodal Intent Recognition via Prototype-Aware Contrastive Alignment and Coarse-to-Fine Dynamic Attention Fusion
por: Huang, Haofeng, et al.
Publicado: (2025)
por: Huang, Haofeng, et al.
Publicado: (2025)
InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees
por: Huang, Chenyu, et al.
Publicado: (2026)
por: Huang, Chenyu, et al.
Publicado: (2026)
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
por: Gupta, Sharut, et al.
Publicado: (2026)
por: Gupta, Sharut, et al.
Publicado: (2026)
Ejemplares similares
-
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
por: Guo, Lixuan, et al.
Publicado: (2026) -
CSRv2: Unlocking Ultra-Sparse Embeddings
por: Guo, Lixuan, et al.
Publicado: (2026) -
Route Experts by Sequence, not by Token
por: Wen, Tiansheng, et al.
Publicado: (2025) -
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
por: Wen, Tiansheng, et al.
Publicado: (2025) -
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
por: Zhang, Qi, et al.
Publicado: (2024)