SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yifan, Su, Zunhai, Hu, Shuhao, Yang, Rui, Wu, Wei, Qian, Yulei, Xie, Yuchen, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
by: Dege, Pengcuo, et al.
Published: (2025)
by: Dege, Pengcuo, et al.
Published: (2025)
Technopolitics at MLA.
by: Kesti, Julie
Published: (1992)
by: Kesti, Julie
Published: (1992)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
by: Li, Guihong, et al.
Published: (2025)
by: Li, Guihong, et al.
Published: (2025)
New Energy for MLA.
by: Shafer, Rita
Published: (1987)
by: Shafer, Rita
Published: (1987)
MLA in San Diego
by: Savage, Noel
Published: (1972)
by: Savage, Noel
Published: (1972)
TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
Whisper-MLA: Reducing GPU Memory Consumption of ASR Models based on MHA2MLA Conversion
by: Zhang, Sen, et al.
Published: (2026)
by: Zhang, Sen, et al.
Published: (2026)
EG-MLA: Embedding-Gated Multi-head Latent Attention for Scalable and Efficient LLMs
by: Cai, Zhengge, et al.
Published: (2025)
by: Cai, Zhengge, et al.
Published: (2025)
MLA Handbook for Writers of Research Papers. Fourth Edition.
by: Gibaldi, Joseph
Published: (1995)
by: Gibaldi, Joseph
Published: (1995)
MLA Handbook for Writers of Research Papers, Theses, and Dissertations.
by: Gibaldi, Joseph, et al.
Published: (1977)
by: Gibaldi, Joseph, et al.
Published: (1977)
Navigating the MLA Bibliography: Performance across Vendor Platforms
by: Soules, Aline, et al.
Published: (2009)
by: Soules, Aline, et al.
Published: (2009)
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
by: Yang, Xiao, et al.
Published: (2025)
by: Yang, Xiao, et al.
Published: (2025)
Geochemistry of Mlanga (MLA) peat core from Kitulu, Tanzania
by: Marchant, Robert
Published: (2021)
by: Marchant, Robert
Published: (2021)
TransMLA: Multi-Head Latent Attention Is All You Need
by: Meng, Fanxu, et al.
Published: (2025)
by: Meng, Fanxu, et al.
Published: (2025)
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
Calibrated ages of Mlanga (MLA) peat core from Kitulu, Tanzania
by: Marchant, Robert
Published: (2021)
by: Marchant, Robert
Published: (2021)
Age determination of Mlanga (MLA) peat core from Kitulu, Tanzania
by: Marchant, Robert
Published: (2021)
by: Marchant, Robert
Published: (2021)
Report of the "MLA Bibliography" Scope and Overlap Committee, an ACRL Ad Hoc Committee.
Published: (1997)
Published: (1997)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
by: Yesiltepe, Hidir, et al.
Published: (2026)
by: Yesiltepe, Hidir, et al.
Published: (2026)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
by: Qian, Yulei, et al.
Published: (2024)
by: Qian, Yulei, et al.
Published: (2024)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
by: Liu, Zhuoyang, et al.
Published: (2025)
by: Liu, Zhuoyang, et al.
Published: (2025)
Amazon AWS Certified Machine Learning Engineer - Associate MLA-C01 PDF
by: Certification Exam
Published: (2026)
by: Certification Exam
Published: (2026)
T-MLA: A targeted multiscale log-exponential attack framework for neural image compression
by: Kalmykov, Nikolay I., et al.
Published: (2025)
by: Kalmykov, Nikolay I., et al.
Published: (2025)
Evaluating the "MLA International Bibliography" for Social Science Content: What Information Can Be Found?
by: Armento, Greg
Published: (1999)
by: Armento, Greg
Published: (1999)
FP8-Flow-MoE: A Casting-Free FP8 Recipe without Double Quantization Error
by: Wang, Fengjuan, et al.
Published: (2025)
by: Wang, Fengjuan, et al.
Published: (2025)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
by: Li, Jonathan, et al.
Published: (2025)
by: Li, Jonathan, et al.
Published: (2025)
P‐7.21: Influence of geometric parameters of Undercut‐like profile on MLA light extraction efficiency
by: Linqing Liu, et al.
Published: (2025)
by: Linqing Liu, et al.
Published: (2025)
Pipelined Decoder for Efficient Context-Aware Text Generation
by: Huang, Zixian, et al.
Published: (2025)
by: Huang, Zixian, et al.
Published: (2025)
FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control
by: Sun, Pingwei, et al.
Published: (2026)
by: Sun, Pingwei, et al.
Published: (2026)
Efficient Post-training Quantization with FP8 Formats
by: Shen, Haihao, et al.
Published: (2023)
by: Shen, Haihao, et al.
Published: (2023)
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
by: Xu, Shuhao, et al.
Published: (2026)
by: Xu, Shuhao, et al.
Published: (2026)
Efficient Context Scaling with LongCat ZigZag Attention
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
4‐2: Design of MLA‐based Integrated Projection System Using Radial Basis Function Mapping Method
by: Xilong Dai, et al.
Published: (2024)
by: Xilong Dai, et al.
Published: (2024)
FP8 Quantization: The Power of the Exponent
by: Kuzmin, Andrey, et al.
Published: (2022)
by: Kuzmin, Andrey, et al.
Published: (2022)
Form and Style: Theses, Reports, Term Papers. Eighth Edition. Up-to-Date Information on Chicago, MLA, and APA Documentation.
by: Campbell, William Giles, et al.
Published: (1990)
by: Campbell, William Giles, et al.
Published: (1990)
HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
by: Liu, Wenxuan, et al.
Published: (2024)
by: Liu, Wenxuan, et al.
Published: (2024)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026)
by: Fan, Xiaoran, et al.
Published: (2026)
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
by: Cai, Shaofei, et al.
Published: (2025)
by: Cai, Shaofei, et al.
Published: (2025)
Similar Items
-
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
by: Dege, Pengcuo, et al.
Published: (2025) -
Technopolitics at MLA.
by: Kesti, Julie
Published: (1992) -
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
by: Li, Guihong, et al.
Published: (2025) -
New Energy for MLA.
by: Shafer, Rita
Published: (1987) -
MLA in San Diego
by: Savage, Noel
Published: (1972)