Scaling Attention via Feature Sparsity
Fuente:
arXiv
Salvato in:
| Autori principali: | Xie, Yan, Wen, Tiansheng, Huang, Tangda, Chen, Bo, You, Chenyu, Jegelka, Stefanie, Wang, Yifei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
di: Guo, Lixuan, et al.
Pubblicazione: (2026)
di: Guo, Lixuan, et al.
Pubblicazione: (2026)
CSRv2: Unlocking Ultra-Sparse Embeddings
di: Guo, Lixuan, et al.
Pubblicazione: (2026)
di: Guo, Lixuan, et al.
Pubblicazione: (2026)
Route Experts by Sequence, not by Token
di: Wen, Tiansheng, et al.
Pubblicazione: (2025)
di: Wen, Tiansheng, et al.
Pubblicazione: (2025)
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
di: Wen, Tiansheng, et al.
Pubblicazione: (2025)
di: Wen, Tiansheng, et al.
Pubblicazione: (2025)
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
di: Zhang, Qi, et al.
Pubblicazione: (2024)
di: Zhang, Qi, et al.
Pubblicazione: (2024)
How to Craft Backdoors with Unlabeled Data Alone?
di: Wang, Yifei, et al.
Pubblicazione: (2024)
di: Wang, Yifei, et al.
Pubblicazione: (2024)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
di: Guo, Xiaojun, et al.
Pubblicazione: (2025)
di: Guo, Xiaojun, et al.
Pubblicazione: (2025)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
Geometric and Dynamic Scaling in Deep Transformers
di: Su, Haoran, et al.
Pubblicazione: (2026)
di: Su, Haoran, et al.
Pubblicazione: (2026)
When More is Less: Understanding Chain-of-Thought Length in LLMs
di: Wu, Yuyang, et al.
Pubblicazione: (2025)
di: Wu, Yuyang, et al.
Pubblicazione: (2025)
Contrastive Factor Analysis
di: Duan, Zhibin, et al.
Pubblicazione: (2024)
di: Duan, Zhibin, et al.
Pubblicazione: (2024)
Geometric Algorithms for Neural Combinatorial Optimization with Constraints
di: Karalias, Nikolaos, et al.
Pubblicazione: (2025)
di: Karalias, Nikolaos, et al.
Pubblicazione: (2025)
Learning with Exact Invariances in Polynomial Time
di: Soleymani, Ashkan, et al.
Pubblicazione: (2025)
di: Soleymani, Ashkan, et al.
Pubblicazione: (2025)
Survey on Generalization Theory for Graph Neural Networks
di: Vasileiou, Antonis, et al.
Pubblicazione: (2025)
di: Vasileiou, Antonis, et al.
Pubblicazione: (2025)
An Information Criterion for Controlled Disentanglement of Multimodal Data
di: Wang, Chenyu, et al.
Pubblicazione: (2024)
di: Wang, Chenyu, et al.
Pubblicazione: (2024)
PRISM: Mitigating EHR Data Sparsity via Learning from Missing Feature Calibrated Prototype Patient Representations
di: Zhu, Yinghao, et al.
Pubblicazione: (2023)
di: Zhu, Yinghao, et al.
Pubblicazione: (2023)
HashAttention: Semantic Sparsity for Faster Inference
di: Desai, Aditya, et al.
Pubblicazione: (2024)
di: Desai, Aditya, et al.
Pubblicazione: (2024)
Learning Diffusion Models with Flexible Representation Guidance
di: Wang, Chenyu, et al.
Pubblicazione: (2025)
di: Wang, Chenyu, et al.
Pubblicazione: (2025)
From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation
di: Ren, Yuxin, et al.
Pubblicazione: (2026)
di: Ren, Yuxin, et al.
Pubblicazione: (2026)
Multi-View Graph Feature Propagation for Privacy Preservation and Feature Sparsity
di: Harari, Etzion, et al.
Pubblicazione: (2025)
di: Harari, Etzion, et al.
Pubblicazione: (2025)
Fairness Aware Reward Optimization
di: Choi, Ching Lam, et al.
Pubblicazione: (2026)
di: Choi, Ching Lam, et al.
Pubblicazione: (2026)
The Empirical Impact of Neural Parameter Symmetries, or Lack Thereof
di: Lim, Derek, et al.
Pubblicazione: (2024)
di: Lim, Derek, et al.
Pubblicazione: (2024)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
di: Gatmiry, Khashayar, et al.
Pubblicazione: (2024)
di: Gatmiry, Khashayar, et al.
Pubblicazione: (2024)
On the Stability of Expressive Positional Encodings for Graphs
di: Huang, Yinan, et al.
Pubblicazione: (2023)
di: Huang, Yinan, et al.
Pubblicazione: (2023)
Understanding the Role of Equivariance in Self-supervised Learning
di: Wang, Yifei, et al.
Pubblicazione: (2024)
di: Wang, Yifei, et al.
Pubblicazione: (2024)
Enhancing Imbalanced Node Classification via Curriculum-Guided Feature Learning and Three-Stage Attention Network
di: Fofanah, Abdul Joseph, et al.
Pubblicazione: (2026)
di: Fofanah, Abdul Joseph, et al.
Pubblicazione: (2026)
A Non-negative VAE:the Generalized Gamma Belief Network
di: Duan, Zhibin, et al.
Pubblicazione: (2024)
di: Duan, Zhibin, et al.
Pubblicazione: (2024)
Learning Linear Attention in Polynomial Time
di: Yau, Morris, et al.
Pubblicazione: (2024)
di: Yau, Morris, et al.
Pubblicazione: (2024)
Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning
di: Ma, Guozheng, et al.
Pubblicazione: (2025)
di: Ma, Guozheng, et al.
Pubblicazione: (2025)
Empty SPACE: Cross-Attention Sparsity for Concept Erasure in Diffusion Models
di: Novello, Nicola, et al.
Pubblicazione: (2026)
di: Novello, Nicola, et al.
Pubblicazione: (2026)
FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
di: Gao, Yifei, et al.
Pubblicazione: (2025)
di: Gao, Yifei, et al.
Pubblicazione: (2025)
ScenGAN: Attention-Intensive Generative Model for Uncertainty-Aware Renewable Scenario Forecasting
di: Wu, Yifei, et al.
Pubblicazione: (2025)
di: Wu, Yifei, et al.
Pubblicazione: (2025)
Post-Training Sparse Attention with Double Sparsity
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity
di: Kim, Kwanyoung, et al.
Pubblicazione: (2025)
di: Kim, Kwanyoung, et al.
Pubblicazione: (2025)
Improving MLLM Training Efficiency via Stage-Aware Sparsity
di: Shi, Kean, et al.
Pubblicazione: (2025)
di: Shi, Kean, et al.
Pubblicazione: (2025)
Improving Decision Sparsity
di: Sun, Yiyang, et al.
Pubblicazione: (2024)
di: Sun, Yiyang, et al.
Pubblicazione: (2024)
MVCL-DAF++: Enhancing Multimodal Intent Recognition via Prototype-Aware Contrastive Alignment and Coarse-to-Fine Dynamic Attention Fusion
di: Huang, Haofeng, et al.
Pubblicazione: (2025)
di: Huang, Haofeng, et al.
Pubblicazione: (2025)
InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees
di: Huang, Chenyu, et al.
Pubblicazione: (2026)
di: Huang, Chenyu, et al.
Pubblicazione: (2026)
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
di: Gupta, Sharut, et al.
Pubblicazione: (2026)
di: Gupta, Sharut, et al.
Pubblicazione: (2026)
Documenti analoghi
-
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
di: Guo, Lixuan, et al.
Pubblicazione: (2026) -
CSRv2: Unlocking Ultra-Sparse Embeddings
di: Guo, Lixuan, et al.
Pubblicazione: (2026) -
Route Experts by Sequence, not by Token
di: Wen, Tiansheng, et al.
Pubblicazione: (2025) -
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
di: Wen, Tiansheng, et al.
Pubblicazione: (2025) -
Beyond Interpretability: The Gains of Feature Monosemanticity on Model Robustness
di: Zhang, Qi, et al.
Pubblicazione: (2024)