Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ji, Tao, Guo, Bin, Wu, Yuanbin, Guo, Qipeng, Shen, Lixing, Chen, Zhan, Qiu, Xipeng, Zhang, Qi, Gui, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026)
by: Fan, Xiaoran, et al.
Published: (2026)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
A Comparison of DeepSeek and Other LLMs
by: Gao, Tianchen, et al.
Published: (2025)
by: Gao, Tianchen, et al.
Published: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
by: Lv, Kai, et al.
Published: (2025)
by: Lv, Kai, et al.
Published: (2025)
Can LLMs Assist Computer Education? an Empirical Case Study of DeepSeek
by: Xiao, Dongfu, et al.
Published: (2025)
by: Xiao, Dongfu, et al.
Published: (2025)
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
by: Xin, Huajian, et al.
Published: (2024)
by: Xin, Huajian, et al.
Published: (2024)
Identifying Semantic Induction Heads to Understand In-Context Learning
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
DeepSeek-V3 Technical Report
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
An evaluation of DeepSeek Models in Biomedical Natural Language Processing
by: Zhan, Zaifu, et al.
Published: (2025)
by: Zhan, Zaifu, et al.
Published: (2025)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
Research on Collaborative Governance of AIGC Applications in the DeepSeek Era
by: Shengli Deng, et al.
Published: (2025)
by: Shengli Deng, et al.
Published: (2025)
From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Medical Reasoning in LLMs: An In-Depth Analysis of DeepSeek R1
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
by: DeepSeek-AI, et al.
Published: (2025)
by: DeepSeek-AI, et al.
Published: (2025)
DeepSeek-VL: Towards Real-World Vision-Language Understanding
by: Lu, Haoyu, et al.
Published: (2024)
by: Lu, Haoyu, et al.
Published: (2024)
DeepSeek-Inspired Exploration of RL-based LLMs and Synergy with Wireless Networks: A Survey
by: Qiao, Yu, et al.
Published: (2025)
by: Qiao, Yu, et al.
Published: (2025)
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
by: Zhang, Wenjing, et al.
Published: (2025)
by: Zhang, Wenjing, et al.
Published: (2025)
DeepSeek-OCR: Contexts Optical Compression
by: Wei, Haoran, et al.
Published: (2025)
by: Wei, Haoran, et al.
Published: (2025)
A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice
by: Bai, Yaowei, et al.
Published: (2025)
by: Bai, Yaowei, et al.
Published: (2025)
Quantitative Analysis of Performance Drop in DeepSeek Model Quantization
by: Zhao, Enbo, et al.
Published: (2025)
by: Zhao, Enbo, et al.
Published: (2025)
DeepSeek in Education: Exploring the Transformative Potential of AI‐Driven Educational Intelligence
by: Jian Liao, et al.
Published: (2025)
by: Jian Liao, et al.
Published: (2025)
Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
by: Qiu, Peiran, et al.
Published: (2025)
by: Qiu, Peiran, et al.
Published: (2025)
Mixture of Tunable Experts -- Behavior Modification of DeepSeek-R1 at Inference Time
by: Dahlke, Robert, et al.
Published: (2025)
by: Dahlke, Robert, et al.
Published: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
by: Liu, Xiaoran, et al.
Published: (2025)
by: Liu, Xiaoran, et al.
Published: (2025)
DeepSeek reshaping healthcare in China's tertiary hospitals
by: Chen, Jishizhan, et al.
Published: (2025)
by: Chen, Jishizhan, et al.
Published: (2025)
Memory Analysis on the Training Course of DeepSeek Models
by: Zhang, Ping, et al.
Published: (2025)
by: Zhang, Ping, et al.
Published: (2025)
Safety Evaluation of DeepSeek Models in Chinese Contexts
by: Zhang, Wenjing, et al.
Published: (2025)
by: Zhang, Wenjing, et al.
Published: (2025)
DeepSeek-OCR 2: Visual Causal Flow
by: Wei, Haoran, et al.
Published: (2026)
by: Wei, Haoran, et al.
Published: (2026)
AdaLomo: Low-memory Optimization with Adaptive Learning Rate
by: Lv, Kai, et al.
Published: (2023)
by: Lv, Kai, et al.
Published: (2023)
How to Set the Learning Rate for Large-Scale Pre-training?
by: Zhou, Yunhua, et al.
Published: (2026)
by: Zhou, Yunhua, et al.
Published: (2026)
Large Language Models for Transforming Healthcare: A Perspective on DeepSeek‐R1
by: Jinsong Zhou, et al.
Published: (2025)
by: Jinsong Zhou, et al.
Published: (2025)
ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
by: Chen, Xinhang, et al.
Published: (2025)
by: Chen, Xinhang, et al.
Published: (2025)
LongWanjuan: Towards Systematic Measurement for Long Text Quality
by: Lv, Kai, et al.
Published: (2024)
by: Lv, Kai, et al.
Published: (2024)
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces
by: Shmidman, Shaltiel, et al.
Published: (2025)
by: Shmidman, Shaltiel, et al.
Published: (2025)
CritiQ: Mining Data Quality Criteria from Human Preferences
by: Guo, Honglin, et al.
Published: (2025)
by: Guo, Honglin, et al.
Published: (2025)
Similar Items
-
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026) -
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025) -
A Comparison of DeepSeek and Other LLMs
by: Gao, Tianchen, et al.
Published: (2025) -
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
by: Ying, Zonghao, et al.
Published: (2025) -
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
by: Lv, Kai, et al.
Published: (2025)