DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Younjoo, Dan, Seungkyun, Lee, Junghoo, Park, Jaiyoung, Ahn, Jung Ho |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
An Inquiry into Datacenter TCO for LLM Inference with FP8
di: Kim, Jiwoo, et al.
Pubblicazione: (2025)
di: Kim, Jiwoo, et al.
Pubblicazione: (2025)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
di: Dong, Ximing, et al.
Pubblicazione: (2026)
di: Dong, Ximing, et al.
Pubblicazione: (2026)
Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
di: Moslem, Yasmin, et al.
Pubblicazione: (2026)
di: Moslem, Yasmin, et al.
Pubblicazione: (2026)
Performance Characterization of Expert Router for Scalable LLM Inference
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024)
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
di: Sunesh, Aman, et al.
Pubblicazione: (2026)
di: Sunesh, Aman, et al.
Pubblicazione: (2026)
Iterative Layer Pruning for Efficient Translation Inference
di: Moslem, Yasmin, et al.
Pubblicazione: (2025)
di: Moslem, Yasmin, et al.
Pubblicazione: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
di: Zheng, Wenhao, et al.
Pubblicazione: (2025)
di: Zheng, Wenhao, et al.
Pubblicazione: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
di: Yang, Shang, et al.
Pubblicazione: (2025)
di: Yang, Shang, et al.
Pubblicazione: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
di: Chen, Feiyang, et al.
Pubblicazione: (2025)
di: Chen, Feiyang, et al.
Pubblicazione: (2025)
ISO: Overlap of Computation and Communication within Seqenence For LLM Inference
di: Xiao, Bin, et al.
Pubblicazione: (2024)
di: Xiao, Bin, et al.
Pubblicazione: (2024)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
di: Ding, Jiabiao, et al.
Pubblicazione: (2026)
di: Ding, Jiabiao, et al.
Pubblicazione: (2026)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
di: Ma, Ming, et al.
Pubblicazione: (2025)
di: Ma, Ming, et al.
Pubblicazione: (2025)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
di: Ray, Kaustabha, et al.
Pubblicazione: (2025)
di: Ray, Kaustabha, et al.
Pubblicazione: (2025)
Cheddar: A Swift Fully Homomorphic Encryption Library Designed for GPU Architectures
di: Choi, Wonseok, et al.
Pubblicazione: (2024)
di: Choi, Wonseok, et al.
Pubblicazione: (2024)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
di: Lin, Yujun, et al.
Pubblicazione: (2024)
di: Lin, Yujun, et al.
Pubblicazione: (2024)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
di: Dong, Juechu, et al.
Pubblicazione: (2024)
di: Dong, Juechu, et al.
Pubblicazione: (2024)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
di: Lin, Mao, et al.
Pubblicazione: (2026)
di: Lin, Mao, et al.
Pubblicazione: (2026)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
Block Sparse Flash Attention
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
di: Jeong, Bodon, et al.
Pubblicazione: (2026)
di: Jeong, Bodon, et al.
Pubblicazione: (2026)
Model Compression and Efficient Inference for Large Language Models: A Survey
di: Wang, Wenxiao, et al.
Pubblicazione: (2024)
di: Wang, Wenxiao, et al.
Pubblicazione: (2024)
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
di: Jiang, Jevin, et al.
Pubblicazione: (2026)
di: Jiang, Jevin, et al.
Pubblicazione: (2026)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
di: Shin, Jiho, et al.
Pubblicazione: (2024)
di: Shin, Jiho, et al.
Pubblicazione: (2024)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
di: Gupta, Ahan, et al.
Pubblicazione: (2023)
di: Gupta, Ahan, et al.
Pubblicazione: (2023)
Layer Importance and Hallucination Analysis in Large Language Models via Enhanced Activation Variance-Sparsity
di: Song, Zichen, et al.
Pubblicazione: (2024)
di: Song, Zichen, et al.
Pubblicazione: (2024)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
di: Israel, Daniel, et al.
Pubblicazione: (2025)
di: Israel, Daniel, et al.
Pubblicazione: (2025)
Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
di: Yun, Sungmin, et al.
Pubblicazione: (2025)
di: Yun, Sungmin, et al.
Pubblicazione: (2025)
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
di: Zhang, Qizheng, et al.
Pubblicazione: (2025)
di: Zhang, Qizheng, et al.
Pubblicazione: (2025)
SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference
di: Cavagna, Hiari Pizzini, et al.
Pubblicazione: (2026)
di: Cavagna, Hiari Pizzini, et al.
Pubblicazione: (2026)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
di: Merouani, Massinissa, et al.
Pubblicazione: (2025)
di: Merouani, Massinissa, et al.
Pubblicazione: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
di: An, Zihao, et al.
Pubblicazione: (2025)
di: An, Zihao, et al.
Pubblicazione: (2025)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
di: Chu, Kexin, et al.
Pubblicazione: (2026)
di: Chu, Kexin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
di: Lee, Gunjun, et al.
Pubblicazione: (2025) -
An Inquiry into Datacenter TCO for LLM Inference with FP8
di: Kim, Jiwoo, et al.
Pubblicazione: (2025) -
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
di: Dong, Ximing, et al.
Pubblicazione: (2026) -
Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
di: Moslem, Yasmin, et al.
Pubblicazione: (2026) -
Performance Characterization of Expert Router for Scalable LLM Inference
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024)