Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Minseo, Hooper, Coleman, Tomar, Aditya, Xu, Chenfeng, Farajtabar, Mehrdad, Mahoney, Michael W., Keutzer, Kurt, Gholami, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Residual Context Diffusion Language Models
by: Hu, Yuezhou, et al.
Published: (2026)
by: Hu, Yuezhou, et al.
Published: (2026)
CDLM: Consistency Diffusion Language Models For Faster Sampling
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
AI and Memory Wall
by: Gholami, Amir, et al.
Published: (2024)
by: Gholami, Amir, et al.
Published: (2024)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)
by: Tiwari, Rishabh, et al.
Published: (2025)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
by: Maheswaran, Monishwaran, et al.
Published: (2025)
by: Maheswaran, Monishwaran, et al.
Published: (2025)
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
SqueezeLLM: Dense-and-Sparse Quantization
by: Kim, Sehoon, et al.
Published: (2023)
by: Kim, Sehoon, et al.
Published: (2023)
Multipole Attention for Efficient Long Context Reasoning
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
Squeezed Attention: Accelerating Long Context Length LLM Inference
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
SPEED: Speculative Pipelined Execution for Efficient Decoding
by: Hooper, Coleman, et al.
Published: (2023)
by: Hooper, Coleman, et al.
Published: (2023)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
by: Tiwari, Rishabh, et al.
Published: (2026)
by: Tiwari, Rishabh, et al.
Published: (2026)
ETS: Efficient Tree Search for Inference-Time Scaling
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
by: Subramanian, Shashank, et al.
Published: (2023)
by: Subramanian, Shashank, et al.
Published: (2023)
Characterizing Prompt Compression Methods for Long Context Inference
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
by: Hooper, Coleman, et al.
Published: (2026)
by: Hooper, Coleman, et al.
Published: (2026)
Learned Best-Effort LLM Serving
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
An LLM Compiler for Parallel Function Calling
by: Kim, Sehoon, et al.
Published: (2023)
by: Kim, Sehoon, et al.
Published: (2023)
TinyAgent: Function Calling at the Edge
by: Erdogan, Lutfi Eren, et al.
Published: (2024)
by: Erdogan, Lutfi Eren, et al.
Published: (2024)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
Agentic Test-Time Scaling for WebAgents
by: Lee, Nicholas, et al.
Published: (2026)
by: Lee, Nicholas, et al.
Published: (2026)
SciML Agents: Write the Solver, Not the Solution
by: Gaonkar, Saarth, et al.
Published: (2025)
by: Gaonkar, Saarth, et al.
Published: (2025)
Immiscible Diffusion: Accelerating Diffusion Training with Noise Assignment
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
by: Lee, Nicholas, et al.
Published: (2024)
by: Lee, Nicholas, et al.
Published: (2024)
Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility
by: Li, Yiheng, et al.
Published: (2025)
by: Li, Yiheng, et al.
Published: (2025)
A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision
by: Peng, Chensheng, et al.
Published: (2024)
by: Peng, Chensheng, et al.
Published: (2024)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
TIDE: Every Layer Knows the Token Beneath the Context
by: Jaiswal, Ajay, et al.
Published: (2026)
by: Jaiswal, Ajay, et al.
Published: (2026)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
by: Samragh, Mohammad, et al.
Published: (2025)
by: Samragh, Mohammad, et al.
Published: (2025)
Efficient and Scalable Estimation of Tool Representations in Vector Space
by: Moon, Suhong, et al.
Published: (2024)
by: Moon, Suhong, et al.
Published: (2024)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
by: Nehrdich, Sebastian, et al.
Published: (2026)
by: Nehrdich, Sebastian, et al.
Published: (2026)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
by: Alizadeh, Keivan, et al.
Published: (2026)
by: Alizadeh, Keivan, et al.
Published: (2026)
AMP: Autoregressive Motion Prediction Revisited with Next Token Prediction for Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2024)
by: Jia, Xiaosong, et al.
Published: (2024)
NEP: Autoregressive Image Editing via Next Editing Token Prediction
by: Wu, Huimin, et al.
Published: (2025)
by: Wu, Huimin, et al.
Published: (2025)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token Optimization
by: Kim, Minseo, et al.
Published: (2026)
by: Kim, Minseo, et al.
Published: (2026)
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
by: Liang, Feng, et al.
Published: (2024)
by: Liang, Feng, et al.
Published: (2024)
Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
by: Ghasemi, Narges, et al.
Published: (2025)
by: Ghasemi, Narges, et al.
Published: (2025)
Similar Items
-
Residual Context Diffusion Language Models
by: Hu, Yuezhou, et al.
Published: (2026) -
CDLM: Consistency Diffusion Language Models For Faster Sampling
by: Kim, Minseo, et al.
Published: (2025) -
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
by: Xi, Haocheng, et al.
Published: (2026) -
AI and Memory Wall
by: Gholami, Amir, et al.
Published: (2024) -
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)