Multi-matrix Factorization Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Jingcheng, Li, Houyi, Zhang, Yinmin, Wang, Zili, Zhou, Shuigeng, Zhang, Xiangyu, Shum, Heung-Yeung, Jiang, Daxin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
by: Hu, Jingcheng, et al.
Published: (2025)
by: Hu, Jingcheng, et al.
Published: (2025)
Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models
by: Li, Houyi, et al.
Published: (2025)
by: Li, Houyi, et al.
Published: (2025)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
by: Li, Houyi, et al.
Published: (2025)
by: Li, Houyi, et al.
Published: (2025)
NoiseAR: AutoRegressing Initial Noise Prior for Diffusion Models
by: Li, Zeming, et al.
Published: (2025)
by: Li, Zeming, et al.
Published: (2025)
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
by: Hu, Jingcheng, et al.
Published: (2026)
by: Hu, Jingcheng, et al.
Published: (2026)
Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
by: Li, Houyi, et al.
Published: (2025)
by: Li, Houyi, et al.
Published: (2025)
Optimizing RLHF Training for Large Language Models with Stage Fusion
by: Zhong, Yinmin, et al.
Published: (2024)
by: Zhong, Yinmin, et al.
Published: (2024)
Selective Attention: Enhancing Transformer through Principled Context Control
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
by: Wang, Dingzirui, et al.
Published: (2025)
by: Wang, Dingzirui, et al.
Published: (2025)
Attention Smoothing Is All You Need For Unlearning
by: Zade, Saleh Zare, et al.
Published: (2026)
by: Zade, Saleh Zare, et al.
Published: (2026)
MULTIVERSE: Exposing Large Language Model Alignment Problems in Diverse Worlds
by: Jin, Xiaolong, et al.
Published: (2024)
by: Jin, Xiaolong, et al.
Published: (2024)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
by: Zhong, Yinmin, et al.
Published: (2025)
by: Zhong, Yinmin, et al.
Published: (2025)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
by: Qiu, Quantong, et al.
Published: (2026)
by: Qiu, Quantong, et al.
Published: (2026)
HiCI: Hierarchical Construction-Integration for Long-Context Attention
by: Zeng, Xiangyu, et al.
Published: (2026)
by: Zeng, Xiangyu, et al.
Published: (2026)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
by: Fu, Zichuan, et al.
Published: (2026)
by: Fu, Zichuan, et al.
Published: (2026)
Improving Transformers with Dynamically Composable Multi-Head Attention
by: Xiao, Da, et al.
Published: (2024)
by: Xiao, Da, et al.
Published: (2024)
TensorLLM: Tensorising Multi-Head Attention for Enhanced Reasoning and Compression in LLMs
by: Gu, Yuxuan, et al.
Published: (2025)
by: Gu, Yuxuan, et al.
Published: (2025)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
by: Chen, Lida, et al.
Published: (2025)
by: Chen, Lida, et al.
Published: (2025)
A Closer Look into Mixture-of-Experts in Large Language Models
by: Lo, Ka Man, et al.
Published: (2024)
by: Lo, Ka Man, et al.
Published: (2024)
Kimi Linear: An Expressive, Efficient Attention Architecture
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
by: Hu, Xiao, et al.
Published: (2025)
by: Hu, Xiao, et al.
Published: (2025)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
by: Ma, Qingsen, et al.
Published: (2026)
by: Ma, Qingsen, et al.
Published: (2026)
Gradient-guided Attention Map Editing: Towards Efficient Contextual Hallucination Mitigation
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
MultiGPrompt for Multi-Task Pre-Training and Prompting on Graphs
by: Yu, Xingtong, et al.
Published: (2023)
by: Yu, Xingtong, et al.
Published: (2023)
Sliding Window Attention Training for Efficient Large Language Models
by: Fu, Zichuan, et al.
Published: (2025)
by: Fu, Zichuan, et al.
Published: (2025)
Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity
by: Yu, Lei, et al.
Published: (2024)
by: Yu, Lei, et al.
Published: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models
by: Huang, Yanwen, et al.
Published: (2025)
by: Huang, Yanwen, et al.
Published: (2025)
Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
by: Gao, Bin, et al.
Published: (2024)
by: Gao, Bin, et al.
Published: (2024)
CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
by: Li, Xiaoya, et al.
Published: (2025)
by: Li, Xiaoya, et al.
Published: (2025)
Stop Overvaluing Multi-Agent Debate -- We Must Rethink Evaluation and Embrace Model Heterogeneity
by: Zhang, Hangfan, et al.
Published: (2025)
by: Zhang, Hangfan, et al.
Published: (2025)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
by: Zhang, Haoyue, et al.
Published: (2025)
by: Zhang, Haoyue, et al.
Published: (2025)
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
by: Jin, Chao, et al.
Published: (2024)
by: Jin, Chao, et al.
Published: (2024)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
by: Zhuang, Xialie, et al.
Published: (2025)
by: Zhuang, Xialie, et al.
Published: (2025)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
Reinforcement Learning for Tool-Integrated Interleaved Thinking towards Cross-Domain Generalization
by: Chen, Zhengyu, et al.
Published: (2025)
by: Chen, Zhengyu, et al.
Published: (2025)
Similar Items
-
Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
by: Hu, Jingcheng, et al.
Published: (2025) -
Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models
by: Li, Houyi, et al.
Published: (2025) -
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
by: Li, Houyi, et al.
Published: (2025) -
NoiseAR: AutoRegressing Initial Noise Prior for Diffusion Models
by: Li, Zeming, et al.
Published: (2025) -
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
by: Hu, Jingcheng, et al.
Published: (2026)