LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Chenkai, Jin, Yijie, Li, Jiajun, Tu, Yi, Long, Guoping, Tu, Dandan, Song, Mingcong, Si, Hongjie, Hou, Tianqi, Yan, Junchi, Deng, Zhijie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LEAP: Unlocking dLLM Parallelism via Lookahead Early-Convergence Token Detection
by: Zhang, Haohui, et al.
Published: (2026)
by: Zhang, Haohui, et al.
Published: (2026)
Scaling Speculative Decoding with Lookahead Reasoning
by: Fu, Yichao, et al.
Published: (2025)
by: Fu, Yichao, et al.
Published: (2025)
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
by: Wu, Chengyue, et al.
Published: (2025)
by: Wu, Chengyue, et al.
Published: (2025)
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels
by: Song, Mingcong, et al.
Published: (2024)
by: Song, Mingcong, et al.
Published: (2024)
LightningRL: Breaking the Accuracy-Parallelism Trade-off of Block-wise dLLMs via Reinforcement Learning
by: Hu, Yanzhe, et al.
Published: (2026)
by: Hu, Yanzhe, et al.
Published: (2026)
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
by: Xiao, Zhongyu, et al.
Published: (2026)
by: Xiao, Zhongyu, et al.
Published: (2026)
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
by: Fu, Yichao, et al.
Published: (2024)
by: Fu, Yichao, et al.
Published: (2024)
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
by: Long, Lingkun, et al.
Published: (2026)
by: Long, Lingkun, et al.
Published: (2026)
dLLM: Simple Diffusion Language Modeling
by: Zhou, Zhanhui, et al.
Published: (2026)
by: Zhou, Zhanhui, et al.
Published: (2026)
ES-dLLM: Efficient Inference for Diffusion Large Language Models by Early-Skipping
by: Zhu, Zijian, et al.
Published: (2026)
by: Zhu, Zijian, et al.
Published: (2026)
Mask Tokens as Prophet: Fine-Grained Cache Eviction for Efficient dLLM Inference
by: Huang, Jianuo, et al.
Published: (2025)
by: Huang, Jianuo, et al.
Published: (2025)
Constrained Decoding with Speculative Lookaheads
by: Nakshatri, Nishanth, et al.
Published: (2024)
by: Nakshatri, Nishanth, et al.
Published: (2024)
DeepPrune: Parallel Scaling without Inter-trace Redundancy
by: Tu, Shangqing, et al.
Published: (2025)
by: Tu, Shangqing, et al.
Published: (2025)
Advancing Text-to-3D Generation with Linearized Lookahead Variational Score Distillation
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
by: Lu, Guanxi, et al.
Published: (2025)
by: Lu, Guanxi, et al.
Published: (2025)
Up to 36x Speedup: Mask-based Parallel Inference Paradigm for Key Information Extraction in MLLMs
by: Wang, Xinzhong, et al.
Published: (2026)
by: Wang, Xinzhong, et al.
Published: (2026)
Fast-dLLM v2: Efficient Block-Diffusion LLM
by: Wu, Chengyue, et al.
Published: (2025)
by: Wu, Chengyue, et al.
Published: (2025)
Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
by: Song, Yuerong, et al.
Published: (2025)
by: Song, Yuerong, et al.
Published: (2025)
dParallel: Learnable Parallel Decoding for dLLMs
by: Chen, Zigeng, et al.
Published: (2025)
by: Chen, Zigeng, et al.
Published: (2025)
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
by: Liu, Zhiyuan, et al.
Published: (2025)
by: Liu, Zhiyuan, et al.
Published: (2025)
Elastic-dLLM: Position Preserving Context Compression and Augmentation of Diffusion LLMs
by: Wu, Junyi, et al.
Published: (2026)
by: Wu, Junyi, et al.
Published: (2026)
Scene-Aware Memory Discrimination: Deciding Which Personal Knowledge Stays
by: Zhong, Yijie, et al.
Published: (2026)
by: Zhong, Yijie, et al.
Published: (2026)
In-context KV-Cache Eviction for LLMs via Attention-Gate
by: Zeng, Zihao, et al.
Published: (2024)
by: Zeng, Zihao, et al.
Published: (2024)
Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
by: Tian, Wenjie, et al.
Published: (2026)
by: Tian, Wenjie, et al.
Published: (2026)
Adaptive Prototype Model for Attribute-based Multi-label Few-shot Action Recognition
by: Xiao, Juefeng, et al.
Published: (2025)
by: Xiao, Juefeng, et al.
Published: (2025)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
by: Lee, Sanghyun, et al.
Published: (2025)
by: Lee, Sanghyun, et al.
Published: (2025)
LINOCS: Lookahead Inference of Networked Operators for Continuous Stability
by: Mudrik, Noga, et al.
Published: (2024)
by: Mudrik, Noga, et al.
Published: (2024)
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation
by: Long, Jianzhi, et al.
Published: (2025)
by: Long, Jianzhi, et al.
Published: (2025)
Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models
by: Mamou, Jonathan, et al.
Published: (2024)
by: Mamou, Jonathan, et al.
Published: (2024)
Sample-Efficient "Clustering and Conquer" Procedures for Parallel Large-Scale Ranking and Selection
by: Zhang, Zishi, et al.
Published: (2024)
by: Zhang, Zishi, et al.
Published: (2024)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
by: Zhang, Tianao, et al.
Published: (2025)
by: Zhang, Tianao, et al.
Published: (2025)
DMax: Aggressive Parallel Decoding for dLLMs
by: Chen, Zigeng, et al.
Published: (2026)
by: Chen, Zigeng, et al.
Published: (2026)
Parallel Continuous Chain-of-Thought with Jacobi Iteration
by: Wu, Haoyi, et al.
Published: (2025)
by: Wu, Haoyi, et al.
Published: (2025)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
by: Luo, Jiajun, et al.
Published: (2024)
by: Luo, Jiajun, et al.
Published: (2024)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
by: Lin, Chien-Yu, et al.
Published: (2025)
by: Lin, Chien-Yu, et al.
Published: (2025)
$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction
by: Du, Zhenbang, et al.
Published: (2026)
by: Du, Zhenbang, et al.
Published: (2026)
Accelerating Transformer Inference for Translation via Parallel Decoding
by: Santilli, Andrea, et al.
Published: (2023)
by: Santilli, Andrea, et al.
Published: (2023)
Similar Items
-
LEAP: Unlocking dLLM Parallelism via Lookahead Early-Convergence Token Detection
by: Zhang, Haohui, et al.
Published: (2026) -
Scaling Speculative Decoding with Lookahead Reasoning
by: Fu, Yichao, et al.
Published: (2025) -
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
by: Wu, Chengyue, et al.
Published: (2025) -
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
by: Wang, Xu, et al.
Published: (2025) -
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025)