From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Yuchuan, Liang, Yuchen, Zhang, Shuo, Shu, Yingte, Yang, Guangwen, He, Wei, Fang, Sibo, Guo, Tianyu, Han, Kai, Xu, Chao, Chen, Hanting, Chen, Xinghao, Wang, Yunhe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deferred Commitment Decoding for Diffusion Language Models
by: Shu, Yingte, et al.
Published: (2026)
by: Shu, Yingte, et al.
Published: (2026)
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025)
by: Tian, Yuchuan, et al.
Published: (2025)
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
Nexus: Higher-Order Attention Mechanisms in Transformers
by: Chen, Hanting, et al.
Published: (2025)
by: Chen, Hanting, et al.
Published: (2025)
DiC: Rethinking Conv3x3 Designs in Diffusion Models
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
Learning Quantized Adaptive Conditions for Diffusion Models
by: Liang, Yuchen, et al.
Published: (2024)
by: Liang, Yuchen, et al.
Published: (2024)
DiJiang: Efficient Large Language Models through Compact Kernelization
by: Chen, Hanting, et al.
Published: (2024)
by: Chen, Hanting, et al.
Published: (2024)
Top 10 Open Challenges Steering the Future of Diffusion Language Model and Its Variants
by: Wang, Yunhe, et al.
Published: (2026)
by: Wang, Yunhe, et al.
Published: (2026)
Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
by: Rang, Miao, et al.
Published: (2025)
by: Rang, Miao, et al.
Published: (2025)
ROOT: Robust Orthogonalized Optimizer for Neural Network Training
by: He, Wei, et al.
Published: (2025)
by: He, Wei, et al.
Published: (2025)
Instruct-IPT: All-in-One Image Processing Transformer via Weight Modulation
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
Multiscale Positive-Unlabeled Detection of AI-Generated Texts
by: Tian, Yuchuan, et al.
Published: (2023)
by: Tian, Yuchuan, et al.
Published: (2023)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
LLMs are Not Just Next Token Predictors
by: Downes, Stephen M., et al.
Published: (2024)
by: Downes, Stephen M., et al.
Published: (2024)
Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap
by: Yang, Chun-Hao, et al.
Published: (2025)
by: Yang, Chun-Hao, et al.
Published: (2025)
LogitSpec: Accelerating Retrieval-based Speculative Decoding via Next Next Token Speculation
by: Liu, Tianyu, et al.
Published: (2025)
by: Liu, Tianyu, et al.
Published: (2025)
SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?
by: Yang, Yueyi, et al.
Published: (2026)
by: Yang, Yueyi, et al.
Published: (2026)
Physics-Guided Multimodal Transformers are the Necessary Foundation for the Next Generation of Meteorological Science
by: Han, Jing, et al.
Published: (2025)
by: Han, Jing, et al.
Published: (2025)
VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
by: Nie, Ying, et al.
Published: (2025)
by: Nie, Ying, et al.
Published: (2025)
Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing
by: Rang, Miao, et al.
Published: (2026)
by: Rang, Miao, et al.
Published: (2026)
Object Recognition as Next Token Prediction
by: Yue, Kaiyu, et al.
Published: (2023)
by: Yue, Kaiyu, et al.
Published: (2023)
Multimodal Latent Language Modeling with Next-Token Diffusion
by: Sun, Yutao, et al.
Published: (2024)
by: Sun, Yutao, et al.
Published: (2024)
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
by: Inferix Team, et al.
Published: (2025)
by: Inferix Team, et al.
Published: (2025)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
by: Kilian, Maciej, et al.
Published: (2024)
by: Kilian, Maciej, et al.
Published: (2024)
Where to Move Next: Zero-shot Generalization of LLMs for Next POI Recommendation
by: Feng, Shanshan, et al.
Published: (2024)
by: Feng, Shanshan, et al.
Published: (2024)
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Top-Quark Decay at Next-to-Next-to-Next-to-Leading Order in QCD
by: Chen, Long, et al.
Published: (2023)
by: Chen, Long, et al.
Published: (2023)
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
by: Yang, Shu-wen, et al.
Published: (2025)
by: Yang, Shu-wen, et al.
Published: (2025)
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
by: Meituan LongCat Team, et al.
Published: (2026)
by: Meituan LongCat Team, et al.
Published: (2026)
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking
by: Li, Wenshuo, et al.
Published: (2024)
by: Li, Wenshuo, et al.
Published: (2024)
Synergizing AI and Digital Twins for Next-Generation Network Optimization, Forecasting, and Security
by: Zhang, Zifan, et al.
Published: (2025)
by: Zhang, Zifan, et al.
Published: (2025)
Prot2Token: A Unified Framework for Protein Modeling via Next-Token Prediction
by: Pourmirzaei, Mahdi, et al.
Published: (2025)
by: Pourmirzaei, Mahdi, et al.
Published: (2025)
Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
by: Chen, Liang, et al.
Published: (2024)
by: Chen, Liang, et al.
Published: (2024)
SENTRA: Selected-Next-Token Transformer for LLM Text Detection
by: Plyler, Mitchell, et al.
Published: (2025)
by: Plyler, Mitchell, et al.
Published: (2025)
Next Tokens Denoising for Speech Synthesis
by: Liu, Yanqing, et al.
Published: (2025)
by: Liu, Yanqing, et al.
Published: (2025)
Next-Token Prediction and Regret Minimization
by: Mohri, Mehryar, et al.
Published: (2026)
by: Mohri, Mehryar, et al.
Published: (2026)
Humanoid Locomotion as Next Token Prediction
by: Radosavovic, Ilija, et al.
Published: (2024)
by: Radosavovic, Ilija, et al.
Published: (2024)
In-Context Imitation Learning via Next-Token Prediction
by: Fu, Letian, et al.
Published: (2024)
by: Fu, Letian, et al.
Published: (2024)
PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMs
by: Zhou, Teng, et al.
Published: (2024)
by: Zhou, Teng, et al.
Published: (2024)
Similar Items
-
Deferred Commitment Decoding for Diffusion Language Models
by: Shu, Yingte, et al.
Published: (2026) -
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025) -
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
by: Tian, Yuchuan, et al.
Published: (2024) -
Nexus: Higher-Order Attention Mechanisms in Transformers
by: Chen, Hanting, et al.
Published: (2025) -
DiC: Rethinking Conv3x3 Designs in Diffusion Models
by: Tian, Yuchuan, et al.
Published: (2024)