X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Guihong, Rezagholizadeh, Mehdi, Yang, Mingyu, Appia, Vikram, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zebra-Llama: Towards Extremely Efficient Hybrid Models
by: Yang, Mingyu, et al.
Published: (2025)
by: Yang, Mingyu, et al.
Published: (2025)
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
by: Fashi, Parsa Ashrafi, et al.
Published: (2026)
by: Fashi, Parsa Ashrafi, et al.
Published: (2026)
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
by: Zhang, Yifan, et al.
Published: (2026)
by: Zhang, Yifan, et al.
Published: (2026)
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
by: Haridas, Akash, et al.
Published: (2026)
by: Haridas, Akash, et al.
Published: (2026)
EG-MLA: Embedding-Gated Multi-head Latent Attention for Scalable and Efficient LLMs
by: Cai, Zhengge, et al.
Published: (2025)
by: Cai, Zhengge, et al.
Published: (2025)
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
by: Joshi, Vinay, et al.
Published: (2025)
by: Joshi, Vinay, et al.
Published: (2025)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
by: Dege, Pengcuo, et al.
Published: (2025)
by: Dege, Pengcuo, et al.
Published: (2025)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
by: Younesian, Sharareh, et al.
Published: (2026)
by: Younesian, Sharareh, et al.
Published: (2026)
Technopolitics at MLA.
by: Kesti, Julie
Published: (1992)
by: Kesti, Julie
Published: (1992)
FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
by: Dukler, Yonatan, et al.
Published: (2025)
by: Dukler, Yonatan, et al.
Published: (2025)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
by: Fan, Xiaoran, et al.
Published: (2026)
by: Fan, Xiaoran, et al.
Published: (2026)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
by: Yang, Dongquan, et al.
Published: (2025)
by: Yang, Dongquan, et al.
Published: (2025)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
by: Yesiltepe, Hidir, et al.
Published: (2026)
by: Yesiltepe, Hidir, et al.
Published: (2026)
New Energy for MLA.
by: Shafer, Rita
Published: (1987)
by: Shafer, Rita
Published: (1987)
MLA in San Diego
by: Savage, Noel
Published: (1972)
by: Savage, Noel
Published: (1972)
TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
by: Yüzügüler, Ahmet Caner, et al.
Published: (2025)
ReGLA: Refining Gated Linear Attention
by: Lu, Peng, et al.
Published: (2025)
by: Lu, Peng, et al.
Published: (2025)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
by: Rehg, Isaac
Published: (2024)
by: Rehg, Isaac
Published: (2024)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
by: Tang, Hanlin, et al.
Published: (2024)
by: Tang, Hanlin, et al.
Published: (2024)
KV Shifting Attention Enhances Language Modeling
by: Xu, Mingyu, et al.
Published: (2024)
by: Xu, Mingyu, et al.
Published: (2024)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
by: Mao, Weian, et al.
Published: (2026)
by: Mao, Weian, et al.
Published: (2026)
Whisper-MLA: Reducing GPU Memory Consumption of ASR Models based on MHA2MLA Conversion
by: Zhang, Sen, et al.
Published: (2026)
by: Zhang, Sen, et al.
Published: (2026)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
by: Yang, Qingyue, et al.
Published: (2025)
by: Yang, Qingyue, et al.
Published: (2025)
GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression
by: Goldstein, Daniel, et al.
Published: (2024)
by: Goldstein, Daniel, et al.
Published: (2024)
MLA Handbook for Writers of Research Papers. Fourth Edition.
by: Gibaldi, Joseph
Published: (1995)
by: Gibaldi, Joseph
Published: (1995)
Effectively Compress KV Heads for LLM
by: Yu, Hao, et al.
Published: (2024)
by: Yu, Hao, et al.
Published: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
by: Ji, Shiyu, et al.
Published: (2026)
by: Ji, Shiyu, et al.
Published: (2026)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
by: Liao, Huanxuan, et al.
Published: (2025)
by: Liao, Huanxuan, et al.
Published: (2025)
The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
by: Guan, Bryan, et al.
Published: (2025)
by: Guan, Bryan, et al.
Published: (2025)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
Navigating the MLA Bibliography: Performance across Vendor Platforms
by: Soules, Aline, et al.
Published: (2009)
by: Soules, Aline, et al.
Published: (2009)
SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
by: Li, Mengjie, et al.
Published: (2025)
by: Li, Mengjie, et al.
Published: (2025)
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
TransMLA: Multi-Head Latent Attention Is All You Need
by: Meng, Fanxu, et al.
Published: (2025)
by: Meng, Fanxu, et al.
Published: (2025)
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
by: Ye, Lu, et al.
Published: (2024)
by: Ye, Lu, et al.
Published: (2024)
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity
by: Ma, Da, et al.
Published: (2024)
by: Ma, Da, et al.
Published: (2024)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
Similar Items
-
Zebra-Llama: Towards Extremely Efficient Hybrid Models
by: Yang, Mingyu, et al.
Published: (2025) -
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
by: Fashi, Parsa Ashrafi, et al.
Published: (2026) -
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
by: Zhang, Yifan, et al.
Published: (2026) -
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
by: Haridas, Akash, et al.
Published: (2026) -
EG-MLA: Embedding-Gated Multi-head Latent Attention for Scalable and Efficient LLMs
by: Cai, Zhengge, et al.
Published: (2025)