Fast and Accurate Causal Parallel Decoding using Jacobi Forcing
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Lanxiang, Kou, Siqi, Fu, Yichao, Rajbhandari, Samyam, Rosing, Tajana, He, Yuxiong, Deng, Zhijie, Zhang, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
by: Hu, Lanxiang, et al.
Published: (2024)
by: Hu, Lanxiang, et al.
Published: (2024)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
by: Qiao, Aurick, et al.
Published: (2024)
by: Qiao, Aurick, et al.
Published: (2024)
CLLMs: Consistency Large Language Models
by: Kou, Siqi, et al.
Published: (2024)
by: Kou, Siqi, et al.
Published: (2024)
FastKernels: Benchmarking GPU Kernel Generation in Production
by: Oliaro, Gabriele, et al.
Published: (2026)
by: Oliaro, Gabriele, et al.
Published: (2026)
Scaling Speculative Decoding with Lookahead Reasoning
by: Fu, Yichao, et al.
Published: (2025)
by: Fu, Yichao, et al.
Published: (2025)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
by: Hidayetoglu, Mert, et al.
Published: (2025)
by: Hidayetoglu, Mert, et al.
Published: (2025)
Online Speculative Decoding
by: Liu, Xiaoxuan, et al.
Published: (2023)
by: Liu, Xiaoxuan, et al.
Published: (2023)
Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI
by: Rajbhandari, Samyam, et al.
Published: (2025)
by: Rajbhandari, Samyam, et al.
Published: (2025)
Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
by: Zhao, Daniel, et al.
Published: (2025)
by: Zhao, Daniel, et al.
Published: (2025)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
by: Lee, Jaeseong, et al.
Published: (2025)
by: Lee, Jaeseong, et al.
Published: (2025)
Efficiently Scaling LLM Reasoning with Certaindex
by: Fu, Yichao, et al.
Published: (2024)
by: Fu, Yichao, et al.
Published: (2024)
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
by: Reichman, Benjamin, et al.
Published: (2025)
by: Reichman, Benjamin, et al.
Published: (2025)
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
by: Fu, Yichao, et al.
Published: (2024)
by: Fu, Yichao, et al.
Published: (2024)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
by: Wang, Xuan, et al.
Published: (2024)
by: Wang, Xuan, et al.
Published: (2024)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
by: Lin, Bokai, et al.
Published: (2024)
by: Lin, Bokai, et al.
Published: (2024)
d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
by: Qian, Yu-Yang, et al.
Published: (2026)
by: Qian, Yu-Yang, et al.
Published: (2026)
Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences
by: Bekman, Stas, et al.
Published: (2025)
by: Bekman, Stas, et al.
Published: (2025)
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
by: Su, Zhaoyuan, et al.
Published: (2026)
by: Su, Zhaoyuan, et al.
Published: (2026)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
by: Yang, Lijie, et al.
Published: (2024)
by: Yang, Lijie, et al.
Published: (2024)
MicroHD: An Accuracy-Driven Optimization of Hyperdimensional Computing Algorithms for TinyML systems
by: Ponzina, Flavio, et al.
Published: (2024)
by: Ponzina, Flavio, et al.
Published: (2024)
LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering
by: Gungor, Onat, et al.
Published: (2025)
by: Gungor, Onat, et al.
Published: (2025)
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
by: Wu, Chengyue, et al.
Published: (2025)
by: Wu, Chengyue, et al.
Published: (2025)
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
by: Tian, Ye, et al.
Published: (2025)
by: Tian, Ye, et al.
Published: (2025)
Thinking with Generated Images
by: Chern, Ethan, et al.
Published: (2025)
by: Chern, Ethan, et al.
Published: (2025)
Parallel Continuous Chain-of-Thought with Jacobi Iteration
by: Wu, Haoyi, et al.
Published: (2025)
by: Wu, Haoyi, et al.
Published: (2025)
SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions
by: Yu, Xiaofan, et al.
Published: (2025)
by: Yu, Xiaofan, et al.
Published: (2025)
Fast Chain-of-Thought: A Glance of Future from Parallel Decoding Leads to Answers Faster
by: Zhang, Hongxuan, et al.
Published: (2023)
by: Zhang, Hongxuan, et al.
Published: (2023)
Augmented Weak Distance for Fast and Accurate Bounds Checking
by: Fu, Zhoulai, et al.
Published: (2025)
by: Fu, Zhoulai, et al.
Published: (2025)
lmgame-Bench: How Good are LLMs at Playing Games?
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
FaTRQ: Tiered Residual Quantization for LLM Vector Search in Far-Memory-Aware ANNS Systems
by: Zhang, Tianqi, et al.
Published: (2026)
by: Zhang, Tianqi, et al.
Published: (2026)
SpANNS: Optimizing Approximate Nearest Neighbor Search for Sparse Vectors Using Near Memory Processing
by: Zhang, Tianqi, et al.
Published: (2026)
by: Zhang, Tianqi, et al.
Published: (2026)
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
by: Holmes, Connor, et al.
Published: (2024)
by: Holmes, Connor, et al.
Published: (2024)
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
by: Huang, Jianuo, et al.
Published: (2026)
by: Huang, Jianuo, et al.
Published: (2026)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
dParallel: Learnable Parallel Decoding for dLLMs
by: Chen, Zigeng, et al.
Published: (2025)
by: Chen, Zigeng, et al.
Published: (2025)
PEARL: Parallel Speculative Decoding with Adaptive Draft Length
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
A KL Lens on Quantization: Fast, Forward-Only Sensitivity for Mixed-Precision SSM-Transformer Models
by: Kong, Jason, et al.
Published: (2026)
by: Kong, Jason, et al.
Published: (2026)
Decoding at the Speed of Thought: Harnessing Parallel Decoding of Lexical Units for LLMs
by: Sun, Chenxi, et al.
Published: (2024)
by: Sun, Chenxi, et al.
Published: (2024)
Similar Items
-
TrimLLM: Progressive Layer Dropping for Domain-Specific LLMs
by: Hu, Lanxiang, et al.
Published: (2024) -
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
by: Qiao, Aurick, et al.
Published: (2024) -
CLLMs: Consistency Large Language Models
by: Kou, Siqi, et al.
Published: (2024) -
FastKernels: Benchmarking GPU Kernel Generation in Production
by: Oliaro, Gabriele, et al.
Published: (2026) -
Scaling Speculative Decoding with Lookahead Reasoning
by: Fu, Yichao, et al.
Published: (2025)