Rethinking Token Prediction: Tree-Structured Diffusion Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Zihao, Yang, Haoming, Dong, Juncheng, Tarokh, Vahid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
S2TX: Cross-Attention Multi-Scale State-Space Transformer for Time Series Forecasting
von: Wu, Zihao, et al.
Veröffentlicht: (2025)
von: Wu, Zihao, et al.
Veröffentlicht: (2025)
Boosting In-Context Learning in LLMs Through the Lens of Classical Supervised Learning
von: Gundem, Korel, et al.
Veröffentlicht: (2025)
von: Gundem, Korel, et al.
Veröffentlicht: (2025)
Score-Based Metropolis-Hastings Algorithms
von: Aloui, Ahmed, et al.
Veröffentlicht: (2024)
von: Aloui, Ahmed, et al.
Veröffentlicht: (2024)
Teleportation With Null Space Gradient Projection for Optimization Acceleration
von: Wu, Zihao, et al.
Veröffentlicht: (2025)
von: Wu, Zihao, et al.
Veröffentlicht: (2025)
Parabolic Continual Learning
von: Yang, Haoming, et al.
Veröffentlicht: (2025)
von: Yang, Haoming, et al.
Veröffentlicht: (2025)
Offline Stochastic Optimization of Black-Box Objective Functions
von: Dong, Juncheng, et al.
Veröffentlicht: (2024)
von: Dong, Juncheng, et al.
Veröffentlicht: (2024)
Rethinking Token Reduction for State Space Models
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
von: Zhan, Zheng, et al.
Veröffentlicht: (2024)
ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models
von: Zheng, Kangjie, et al.
Veröffentlicht: (2025)
von: Zheng, Kangjie, et al.
Veröffentlicht: (2025)
Neural McKean-Vlasov Processes: Distributional Dependence in Diffusion Processes
von: Yang, Haoming, et al.
Veröffentlicht: (2024)
von: Yang, Haoming, et al.
Veröffentlicht: (2024)
Parallel Token Prediction for Language Models
von: Draxler, Felix, et al.
Veröffentlicht: (2025)
von: Draxler, Felix, et al.
Veröffentlicht: (2025)
Contextual Text Denoising with Masked Language Models
von: Sun, Yifu, et al.
Veröffentlicht: (2019)
von: Sun, Yifu, et al.
Veröffentlicht: (2019)
Conditional Average Treatment Effect Estimation Under Hidden Confounders
von: Aloui, Ahmed, et al.
Veröffentlicht: (2025)
von: Aloui, Ahmed, et al.
Veröffentlicht: (2025)
Length-MAX Tokenizer for Language Models
von: Dong, Dong, et al.
Veröffentlicht: (2025)
von: Dong, Dong, et al.
Veröffentlicht: (2025)
Multimodal Latent Language Modeling with Next-Token Diffusion
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical Prediction
von: Zhu, Mingcheng, et al.
Veröffentlicht: (2026)
von: Zhu, Mingcheng, et al.
Veröffentlicht: (2026)
Efficient Temporal Tokenization for Mobility Prediction with Large Language Models
von: He, Haoyu, et al.
Veröffentlicht: (2025)
von: He, Haoyu, et al.
Veröffentlicht: (2025)
Just on Time: Token-Level Early Stopping for Diffusion Language Models
von: Kohut, Zahar, et al.
Veröffentlicht: (2026)
von: Kohut, Zahar, et al.
Veröffentlicht: (2026)
Elliptic Loss Regularization
von: Hasan, Ali, et al.
Veröffentlicht: (2025)
von: Hasan, Ali, et al.
Veröffentlicht: (2025)
CATE Estimation With Potential Outcome Imputation From Local Regression
von: Aloui, Ahmed, et al.
Veröffentlicht: (2023)
von: Aloui, Ahmed, et al.
Veröffentlicht: (2023)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
In-Context Reinforcement Learning From Suboptimal Historical Data
von: Dong, Juncheng, et al.
Veröffentlicht: (2026)
von: Dong, Juncheng, et al.
Veröffentlicht: (2026)
Rethinking Machine Unlearning for Large Language Models
von: Liu, Sijia, et al.
Veröffentlicht: (2024)
von: Liu, Sijia, et al.
Veröffentlicht: (2024)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
von: Mao, Yu, et al.
Veröffentlicht: (2025)
von: Mao, Yu, et al.
Veröffentlicht: (2025)
Improving Diffusion Language Model Decoding through Joint Search in Generation Order and Token Space
von: Shen, Yangyi, et al.
Veröffentlicht: (2026)
von: Shen, Yangyi, et al.
Veröffentlicht: (2026)
Tokenization Tradeoffs in Structured EHR Foundation Models
von: Guo, Lin Lawrence, et al.
Veröffentlicht: (2026)
von: Guo, Lin Lawrence, et al.
Veröffentlicht: (2026)
Language Modeling with Learned Meta-Tokens
von: Shah, Alok N., et al.
Veröffentlicht: (2025)
von: Shah, Alok N., et al.
Veröffentlicht: (2025)
A Law of Next-Token Prediction in Large Language Models
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
Differentially Private Next-Token Prediction of Large Language Models
von: Flemings, James, et al.
Veröffentlicht: (2024)
von: Flemings, James, et al.
Veröffentlicht: (2024)
Unsupervised Morphological Tree Tokenizer
von: Zhu, Qingyang, et al.
Veröffentlicht: (2024)
von: Zhu, Qingyang, et al.
Veröffentlicht: (2024)
Stability-Weighted Decoding for Diffusion Language Models
von: Wu, Yue, et al.
Veröffentlicht: (2026)
von: Wu, Yue, et al.
Veröffentlicht: (2026)
Can Perplexity Predict Fine-tuning Performance? An Investigation of Tokenization Effects on Sequential Language Models for Nepali
von: Luitel, Nishant, et al.
Veröffentlicht: (2024)
von: Luitel, Nishant, et al.
Veröffentlicht: (2024)
Rethinking LLM Ensembling from the Perspective of Mixture Models
von: Fu, Jiale, et al.
Veröffentlicht: (2026)
von: Fu, Jiale, et al.
Veröffentlicht: (2026)
CARE: Turning LLMs Into Causal Reasoning Expert
von: Dong, Juncheng, et al.
Veröffentlicht: (2025)
von: Dong, Juncheng, et al.
Veröffentlicht: (2025)
Next Semantic Scale Prediction via Hierarchical Diffusion Language Models
von: Zhou, Cai, et al.
Veröffentlicht: (2025)
von: Zhou, Cai, et al.
Veröffentlicht: (2025)
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models
von: Kongmanee, Jaturong
Veröffentlicht: (2025)
von: Kongmanee, Jaturong
Veröffentlicht: (2025)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
von: Thrampoulidis, Christos
Veröffentlicht: (2024)
von: Thrampoulidis, Christos
Veröffentlicht: (2024)
Diffusion-Based Hypothesis Testing and Change-Point Detection
von: Moushegian, Sean, et al.
Veröffentlicht: (2025)
von: Moushegian, Sean, et al.
Veröffentlicht: (2025)
The Geometry of Tokens in Internal Representations of Large Language Models
von: Viswanathan, Karthik, et al.
Veröffentlicht: (2025)
von: Viswanathan, Karthik, et al.
Veröffentlicht: (2025)
Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling
von: Luo, Xianzhen, et al.
Veröffentlicht: (2024)
von: Luo, Xianzhen, et al.
Veröffentlicht: (2024)
Sequential Diffusion Language Models
von: Liu, Yangzhou, et al.
Veröffentlicht: (2025)
von: Liu, Yangzhou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
S2TX: Cross-Attention Multi-Scale State-Space Transformer for Time Series Forecasting
von: Wu, Zihao, et al.
Veröffentlicht: (2025) -
Boosting In-Context Learning in LLMs Through the Lens of Classical Supervised Learning
von: Gundem, Korel, et al.
Veröffentlicht: (2025) -
Score-Based Metropolis-Hastings Algorithms
von: Aloui, Ahmed, et al.
Veröffentlicht: (2024) -
Teleportation With Null Space Gradient Projection for Optimization Acceleration
von: Wu, Zihao, et al.
Veröffentlicht: (2025) -
Parabolic Continual Learning
von: Yang, Haoming, et al.
Veröffentlicht: (2025)