MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yu, Chen, Runkai, Yi, Sheng, Zhao, Xinda, Li, Xiaohong, Zhang, Jianjin, Sun, Jun, Hu, Chuanrui, Han, Yunyun, Bing, Lidong, Deng, Yafeng, Chen, Tianqiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PIT: A Dynamic Personalized Item Tokenizer for End-to-End Generative Recommendation
by: Wang, Huanjie, et al.
Published: (2026)
by: Wang, Huanjie, et al.
Published: (2026)
Generative Recommender with End-to-End Learnable Item Tokenization
by: Liu, Enze, et al.
Published: (2024)
by: Liu, Enze, et al.
Published: (2024)
LASER: An Efficient Target-Aware Segmented Attention Framework for End-to-End Long Sequence Modeling
by: Lin, Tianhe, et al.
Published: (2026)
by: Lin, Tianhe, et al.
Published: (2026)
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
by: Feng, Pengchao, et al.
Published: (2025)
by: Feng, Pengchao, et al.
Published: (2025)
OpenRAG: Optimizing RAG End-to-End via In-Context Retrieval Learning
by: Zhou, Jiawei, et al.
Published: (2025)
by: Zhou, Jiawei, et al.
Published: (2025)
Differentiable Geometric Indexing for End-to-End Generative Retrieval
by: Wang, Xujing, et al.
Published: (2026)
by: Wang, Xujing, et al.
Published: (2026)
ASI++: Towards Distributionally Balanced End-to-End Generative Retrieval
by: Liu, Yuxuan, et al.
Published: (2024)
by: Liu, Yuxuan, et al.
Published: (2024)
STORE: Semantic Tokenization, Orthogonal Rotation and Efficient Attention for Scaling Up Ranking Models
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
VQL: An End-to-End Context-Aware Vector Quantization Attention for Ultra-Long User Behavior Modeling
by: Li, Kaiyuan, et al.
Published: (2025)
by: Li, Kaiyuan, et al.
Published: (2025)
OneSug: The Unified End-to-End Generative Framework for E-commerce Query Suggestion
by: Guo, Xian, et al.
Published: (2025)
by: Guo, Xian, et al.
Published: (2025)
Sparton: Fast and Memory-Efficient Triton Kernel for Learned Sparse Retrieval
by: Nguyen, Thong, et al.
Published: (2026)
by: Nguyen, Thong, et al.
Published: (2026)
From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs
by: Wu, Yaxiong, et al.
Published: (2025)
by: Wu, Yaxiong, et al.
Published: (2025)
RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
by: Li, Shaobo, et al.
Published: (2026)
by: Li, Shaobo, et al.
Published: (2026)
Beyond the Needle's Illusion: Decoupled Evaluation of Evidence Access and Use under Semantic Interference at 326M-Token Scale
by: Lin, Tianwei, et al.
Published: (2026)
by: Lin, Tianwei, et al.
Published: (2026)
Reasoning to Rank: An End-to-End Solution for Exploiting Large Language Models for Recommendation
by: Zheng, Kehan, et al.
Published: (2026)
by: Zheng, Kehan, et al.
Published: (2026)
Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation
by: Guan, Lin, et al.
Published: (2025)
by: Guan, Lin, et al.
Published: (2025)
Recurrent Preference Memory for Efficient Long-Sequence Generative Recommendation
by: Chen, Yixiao, et al.
Published: (2026)
by: Chen, Yixiao, et al.
Published: (2026)
SARM: LLM-Augmented Semantic Anchor for End-to-End Live-Streaming Ranking
by: Yang, Ruochen, et al.
Published: (2026)
by: Yang, Ruochen, et al.
Published: (2026)
MSN: A Memory-based Sparse Activation Scaling Framework for Large-scale Industrial Recommendation
by: Wu, Shikang, et al.
Published: (2026)
by: Wu, Shikang, et al.
Published: (2026)
Beyong Tokens: Item-aware Attention for LLM-based Recommendation
by: Zhang, Xiaokun, et al.
Published: (2026)
by: Zhang, Xiaokun, et al.
Published: (2026)
Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
by: Hu, Chuanrui, et al.
Published: (2026)
by: Hu, Chuanrui, et al.
Published: (2026)
OneSearch: A Preliminary Exploration of the Unified End-to-End Generative Framework for E-commerce Search
by: Chen, Ben, et al.
Published: (2025)
by: Chen, Ben, et al.
Published: (2025)
UniRank: End-to-End Domain-Specific Reranking of Hybrid Text-Image Candidates
by: Yang, Yupei, et al.
Published: (2026)
by: Yang, Yupei, et al.
Published: (2026)
End-to-End Semantic ID Generation for Generative Advertisement Recommendation
by: Jiang, Jie, et al.
Published: (2026)
by: Jiang, Jie, et al.
Published: (2026)
Towards End-to-End Model-Agnostic Explanations for RAG Systems
by: Sudhi, Viju, et al.
Published: (2025)
by: Sudhi, Viju, et al.
Published: (2025)
MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts
by: Tao, Zhen, et al.
Published: (2026)
by: Tao, Zhen, et al.
Published: (2026)
Memory Assisted LLM for Personalized Recommendation System
by: Chen, Jiarui
Published: (2025)
by: Chen, Jiarui
Published: (2025)
An End-to-End Multi-objective Ensemble Ranking Framework for Video Recommendation
by: He, Tiantian, et al.
Published: (2025)
by: He, Tiantian, et al.
Published: (2025)
EGA-V1: Unifying Online Advertising with End-to-End Learning
by: Qiu, Junyan, et al.
Published: (2025)
by: Qiu, Junyan, et al.
Published: (2025)
Towards End-to-End Alignment of User Satisfaction via Questionnaire in Video Recommendation
by: Li, Na, et al.
Published: (2026)
by: Li, Na, et al.
Published: (2026)
LiCoMemory: Lightweight and Cognitive Agentic Memory for Efficient Long-Term Reasoning
by: Huang, Zhengjun, et al.
Published: (2025)
by: Huang, Zhengjun, et al.
Published: (2025)
Gumbel Reranking: Differentiable End-to-End Reranker Optimization
by: Huang, Siyuan, et al.
Published: (2025)
by: Huang, Siyuan, et al.
Published: (2025)
Large Memory Network for Recommendation
by: Lu, Hui, et al.
Published: (2025)
by: Lu, Hui, et al.
Published: (2025)
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
by: Opsahl-Ong, Krista, et al.
Published: (2026)
by: Opsahl-Ong, Krista, et al.
Published: (2026)
Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
by: Tang, Qiaoyu, et al.
Published: (2024)
by: Tang, Qiaoyu, et al.
Published: (2024)
OneMall: One Architecture, More Scenarios -- End-to-End Generative Recommender Family at Kuaishou E-Commerce
by: Zhang, Kun, et al.
Published: (2026)
by: Zhang, Kun, et al.
Published: (2026)
LEMUR: Large scale End-to-end MUltimodal Recommendation
by: Han, Xintian, et al.
Published: (2025)
by: Han, Xintian, et al.
Published: (2025)
End-to-end training of Multimodal Model and ranking Model
by: Deng, Xiuqi, et al.
Published: (2024)
by: Deng, Xiuqi, et al.
Published: (2024)
Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System
by: Balaka, Muhammad Imam Luthfi, et al.
Published: (2025)
by: Balaka, Muhammad Imam Luthfi, et al.
Published: (2025)
Countering Mainstream Bias via End-to-End Adaptive Local Learning
by: Pan, Jinhao, et al.
Published: (2024)
by: Pan, Jinhao, et al.
Published: (2024)
Similar Items
-
PIT: A Dynamic Personalized Item Tokenizer for End-to-End Generative Recommendation
by: Wang, Huanjie, et al.
Published: (2026) -
Generative Recommender with End-to-End Learnable Item Tokenization
by: Liu, Enze, et al.
Published: (2024) -
LASER: An Efficient Target-Aware Segmented Attention Framework for End-to-End Long Sequence Modeling
by: Lin, Tianhe, et al.
Published: (2026) -
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation
by: Feng, Pengchao, et al.
Published: (2025) -
OpenRAG: Optimizing RAG End-to-End via In-Context Retrieval Learning
by: Zhou, Jiawei, et al.
Published: (2025)