Moving Beyond Next-Token Prediction: Transformers are Context-Sensitive Language Generators
Fuente:
arXiv
Saved in:
| Main Author: | Rhee, Phill Kyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
by: Shao, Chenze, et al.
Published: (2024)
by: Shao, Chenze, et al.
Published: (2024)
Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap
by: Yang, Chun-Hao, et al.
Published: (2025)
by: Yang, Chun-Hao, et al.
Published: (2025)
Alternatives To Next Token Prediction In Text Generation -- A Survey
by: Wyatt, Charlie, et al.
Published: (2025)
by: Wyatt, Charlie, et al.
Published: (2025)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
A Law of Next-Token Prediction in Large Language Models
by: He, Hangfeng, et al.
Published: (2024)
by: He, Hangfeng, et al.
Published: (2024)
Fractal Patterns May Illuminate the Success of Next-Token Prediction
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
by: Alabdulmohsin, Ibrahim, et al.
Published: (2024)
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
by: An, Chenyang, et al.
Published: (2024)
by: An, Chenyang, et al.
Published: (2024)
PICLe: Eliciting Diverse Behaviors from Large Language Models with Persona In-Context Learning
by: Choi, Hyeong Kyu, et al.
Published: (2024)
by: Choi, Hyeong Kyu, et al.
Published: (2024)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
Advancing Pancreatic Cancer Prediction with a Next Visit Token Prediction Head on top of Med-BERT
by: He, Jianping, et al.
Published: (2025)
by: He, Jianping, et al.
Published: (2025)
Text Generation Beyond Discrete Token Sampling
by: Zhuang, Yufan, et al.
Published: (2025)
by: Zhuang, Yufan, et al.
Published: (2025)
LLMs are Not Just Next Token Predictors
by: Downes, Stephen M., et al.
Published: (2024)
by: Downes, Stephen M., et al.
Published: (2024)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
by: Mishra, Shubhra, et al.
Published: (2024)
by: Mishra, Shubhra, et al.
Published: (2024)
Modeling Next-Token Prediction as Left-Nested Intuitionistic Implication
by: Tarau, Paul
Published: (2026)
by: Tarau, Paul
Published: (2026)
Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
by: Hu, Xiang, et al.
Published: (2025)
by: Hu, Xiang, et al.
Published: (2025)
SentenceVAE: Enable Next-sentence Prediction for Large Language Models with Faster Speed, Higher Accuracy and Longer Context
by: An, Hongjun, et al.
Published: (2024)
by: An, Hongjun, et al.
Published: (2024)
From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs
by: Tian, Yuchuan, et al.
Published: (2025)
by: Tian, Yuchuan, et al.
Published: (2025)
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
by: Chun, Jiyun, et al.
Published: (2026)
by: Chun, Jiyun, et al.
Published: (2026)
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models
by: Wei, Sheng-Lun, et al.
Published: (2024)
by: Wei, Sheng-Lun, et al.
Published: (2024)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
by: Jia, Mumin, et al.
Published: (2025)
by: Jia, Mumin, et al.
Published: (2025)
Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations
by: Bochkov, A.
Published: (2025)
by: Bochkov, A.
Published: (2025)
Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models
by: Zhong, Linhao, et al.
Published: (2026)
by: Zhong, Linhao, et al.
Published: (2026)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
by: Wang, Qichao, et al.
Published: (2025)
by: Wang, Qichao, et al.
Published: (2025)
TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG
by: Lu, Pengqian, et al.
Published: (2025)
by: Lu, Pengqian, et al.
Published: (2025)
Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
by: Liu, Peijie, et al.
Published: (2025)
by: Liu, Peijie, et al.
Published: (2025)
Context-level Language Modeling by Learning Predictive Context Embeddings
by: Dai, Beiya, et al.
Published: (2025)
by: Dai, Beiya, et al.
Published: (2025)
Understanding In-Context Learning Beyond Transformers: An Investigation of State Space and Hybrid Architectures
by: Wang, Shenran, et al.
Published: (2025)
by: Wang, Shenran, et al.
Published: (2025)
Pctx: Tokenizing Personalized Context for Generative Recommendation
by: Zhong, Qiyong, et al.
Published: (2025)
by: Zhong, Qiyong, et al.
Published: (2025)
Moving Beyond Review: Applying Language Models to Planning and Translation in Reflection
by: Neshaei, Seyed Parsa, et al.
Published: (2026)
by: Neshaei, Seyed Parsa, et al.
Published: (2026)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
by: Aynetdinov, Ansar, et al.
Published: (2025)
by: Aynetdinov, Ansar, et al.
Published: (2025)
AgentMove: A Large Language Model based Agentic Framework for Zero-shot Next Location Prediction
by: Feng, Jie, et al.
Published: (2024)
by: Feng, Jie, et al.
Published: (2024)
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
by: Liu, Yuliang, et al.
Published: (2026)
by: Liu, Yuliang, et al.
Published: (2026)
SpecFuse: Ensembling Large Language Models via Next-Segment Prediction
by: Lv, Bo, et al.
Published: (2024)
by: Lv, Bo, et al.
Published: (2024)
Evaluating the Sensitivity of LLMs to Prior Context
by: Hankache, Robert, et al.
Published: (2025)
by: Hankache, Robert, et al.
Published: (2025)
Controllable Context Sensitivity and the Knob Behind It
by: Minder, Julian, et al.
Published: (2024)
by: Minder, Julian, et al.
Published: (2024)
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
by: Cho, Gyeongje, et al.
Published: (2025)
by: Cho, Gyeongje, et al.
Published: (2025)
Next Token Perception Score: Analytical Assessment of your LLM Perception Skills
by: Cheng, Yu-Ang, et al.
Published: (2025)
by: Cheng, Yu-Ang, et al.
Published: (2025)
NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training
by: Han, Minglun, et al.
Published: (2024)
by: Han, Minglun, et al.
Published: (2024)
Similar Items
-
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
by: Shao, Chenze, et al.
Published: (2024) -
Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap
by: Yang, Chun-Hao, et al.
Published: (2025) -
Alternatives To Next Token Prediction In Text Generation -- A Survey
by: Wyatt, Charlie, et al.
Published: (2025) -
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025) -
Cautious Next Token Prediction
by: Wang, Yizhou, et al.
Published: (2025)