Guardado en:
| Autores principales: | Yüzügüler, Ahmet Caner, Çelik, Ahmet, Zhuang, Jiawei, Cavigelli, Lukas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2509.21081 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
por: Dege, Pengcuo, et al.
Publicado: (2025)
por: Dege, Pengcuo, et al.
Publicado: (2025)
TransMLA: Multi-Head Latent Attention Is All You Need
por: Meng, Fanxu, et al.
Publicado: (2025)
por: Meng, Fanxu, et al.
Publicado: (2025)
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
por: Ma, Bole, et al.
Publicado: (2026)
por: Ma, Bole, et al.
Publicado: (2026)
Don't be so Stief! Learning KV Cache low-rank approximation over the Stiefel manifold
por: Benfenati, Luca, et al.
Publicado: (2026)
por: Benfenati, Luca, et al.
Publicado: (2026)
SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights
por: Müller, Lorenz K., et al.
Publicado: (2025)
por: Müller, Lorenz K., et al.
Publicado: (2025)
Whisper-MLA: Reducing GPU Memory Consumption of ASR Models based on MHA2MLA Conversion
por: Zhang, Sen, et al.
Publicado: (2026)
por: Zhang, Sen, et al.
Publicado: (2026)
KernelDNA: Dynamic Kernel Sharing via Decoupled Naive Adapters
por: Huang, Haiduo, et al.
Publicado: (2025)
por: Huang, Haiduo, et al.
Publicado: (2025)
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
por: Wang, Xi, et al.
Publicado: (2026)
por: Wang, Xi, et al.
Publicado: (2026)
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
por: Yang, Xiao, et al.
Publicado: (2025)
por: Yang, Xiao, et al.
Publicado: (2025)
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
por: Liu, Zikang, et al.
Publicado: (2025)
por: Liu, Zikang, et al.
Publicado: (2025)
SSSD: Simply-Scalable Speculative Decoding
por: Marzollo, Michele, et al.
Publicado: (2024)
por: Marzollo, Michele, et al.
Publicado: (2024)
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining
por: Zhang, Yifan, et al.
Publicado: (2026)
por: Zhang, Yifan, et al.
Publicado: (2026)
Auto-Unrolled Proximal Gradient Descent: An AutoML Approach to Interpretable Waveform Optimization
por: Kaplan, Ahmet
Publicado: (2026)
por: Kaplan, Ahmet
Publicado: (2026)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
por: Fan, Xiaoran, et al.
Publicado: (2026)
por: Fan, Xiaoran, et al.
Publicado: (2026)
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs
por: Athiwaratkun, Ben, et al.
Publicado: (2024)
por: Athiwaratkun, Ben, et al.
Publicado: (2024)
GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units
por: Bouvier, Maxence, et al.
Publicado: (2025)
por: Bouvier, Maxence, et al.
Publicado: (2025)
KS-Net: Multi-layer network model for determining the rotor type from motor parameters in interior PMSMs
por: Dogan, Kivanc, et al.
Publicado: (2025)
por: Dogan, Kivanc, et al.
Publicado: (2025)
Hierarchical Training of Deep Neural Networks Using Early Exiting
por: Sepehri, Yamin, et al.
Publicado: (2023)
por: Sepehri, Yamin, et al.
Publicado: (2023)
Stable-by-Design Neural Network-Based LPV State-Space Models for System Identification
por: Sertbaş, Ahmet Eren, et al.
Publicado: (2025)
por: Sertbaş, Ahmet Eren, et al.
Publicado: (2025)
Marconi: Prefix Caching for the Era of Hybrid LLMs
por: Pan, Rui, et al.
Publicado: (2024)
por: Pan, Rui, et al.
Publicado: (2024)
The Blind Normalized Stein Variational Gradient Descent-Based Detection for Intelligent Random Access in Cellular IoT
por: Zhu, Xin, et al.
Publicado: (2024)
por: Zhu, Xin, et al.
Publicado: (2024)
ENTIRe-ID: An Extensive and Diverse Dataset for Person Re-Identification
por: Yildiz, Serdar, et al.
Publicado: (2024)
por: Yildiz, Serdar, et al.
Publicado: (2024)
PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
por: Ding, Ruogu, et al.
Publicado: (2025)
por: Ding, Ruogu, et al.
Publicado: (2025)
Multimodal Stock Price Prediction
por: Karadaş, Furkan, et al.
Publicado: (2025)
por: Karadaş, Furkan, et al.
Publicado: (2025)
MediaMind: Revolutionizing Media Monitoring using Agentification
por: Gunduz, Ahmet, et al.
Publicado: (2025)
por: Gunduz, Ahmet, et al.
Publicado: (2025)
Naïve Bayes and Random Forest for Crop Yield Prediction
por: Maazallahi, Abbas, et al.
Publicado: (2024)
por: Maazallahi, Abbas, et al.
Publicado: (2024)
Quantum-Enhanced Parameter-Efficient Learning for Typhoon Trajectory Forecasting
por: Liu, Chen-Yu, et al.
Publicado: (2025)
por: Liu, Chen-Yu, et al.
Publicado: (2025)
Bias-Restrained Prefix Representation Finetuning for Mathematical Reasoning
por: Liang, Sirui, et al.
Publicado: (2025)
por: Liang, Sirui, et al.
Publicado: (2025)
Fast and Effective On-policy Distillation from Reasoning Prefixes
por: Zhang, Dongxu, et al.
Publicado: (2026)
por: Zhang, Dongxu, et al.
Publicado: (2026)
Sparse Prefix Caching for Hybrid and Recurrent LLM Serving
por: Shirokikh, Mikhail, et al.
Publicado: (2026)
por: Shirokikh, Mikhail, et al.
Publicado: (2026)
L-Lipschitz Gershgorin ResNet Network
por: Juston, Marius F. R., et al.
Publicado: (2025)
por: Juston, Marius F. R., et al.
Publicado: (2025)
Q-Sat AI: Machine Learning-Based Decision Support for Data Saturation in Qualitative Studies
por: Tutar, Hasan, et al.
Publicado: (2025)
por: Tutar, Hasan, et al.
Publicado: (2025)
NeuralPrefix: A Zero-shot Sensory Data Imputation Plugin
por: Khamis, Abdelwahed, et al.
Publicado: (2025)
por: Khamis, Abdelwahed, et al.
Publicado: (2025)
A Step Back: Prefix Importance Ratio Stabilizes Policy Optimization
por: Lei, Shiye, et al.
Publicado: (2026)
por: Lei, Shiye, et al.
Publicado: (2026)
1-Lipschitz Network Initialization for Certifiably Robust Classification Applications: A Decay Problem
por: Juston, Marius F. R., et al.
Publicado: (2025)
por: Juston, Marius F. R., et al.
Publicado: (2025)
Towards Infinite-Long Prefix in Transformer
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
Learning Physics Informed Neural ODEs With Partial Measurements
por: Ghanem, Paul, et al.
Publicado: (2024)
por: Ghanem, Paul, et al.
Publicado: (2024)
Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs
por: Ashok, Arjun, et al.
Publicado: (2025)
por: Ashok, Arjun, et al.
Publicado: (2025)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
por: Yesiltepe, Hidir, et al.
Publicado: (2026)
por: Yesiltepe, Hidir, et al.
Publicado: (2026)
Ejemplares similares
-
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025) -
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
por: Dege, Pengcuo, et al.
Publicado: (2025) -
TransMLA: Multi-Head Latent Attention Is All You Need
por: Meng, Fanxu, et al.
Publicado: (2025) -
Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
por: Ma, Bole, et al.
Publicado: (2026) -
Don't be so Stief! Learning KV Cache low-rank approximation over the Stiefel manifold
por: Benfenati, Luca, et al.
Publicado: (2026)