Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Yize, Thrampoulidis, Christos |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
di: Zhao, Yize, et al.
Pubblicazione: (2024)
di: Zhao, Yize, et al.
Pubblicazione: (2024)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
di: Thrampoulidis, Christos
Pubblicazione: (2024)
di: Thrampoulidis, Christos
Pubblicazione: (2024)
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
di: Deng, Wenlong, et al.
Pubblicazione: (2024)
di: Deng, Wenlong, et al.
Pubblicazione: (2024)
NITP: Next Implicit Token Prediction for LLM Pre-training
di: Zhang, Xiangdong, et al.
Pubblicazione: (2026)
di: Zhang, Xiangdong, et al.
Pubblicazione: (2026)
Why Loss Re-weighting Works If You Stop Early: Training Dynamics of Unconstrained Features
di: Zhao, Yize, et al.
Pubblicazione: (2026)
di: Zhao, Yize, et al.
Pubblicazione: (2026)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
di: Deora, Puneesh, et al.
Pubblicazione: (2025)
di: Deora, Puneesh, et al.
Pubblicazione: (2025)
Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
di: Vakilian, Vala, et al.
Pubblicazione: (2025)
di: Vakilian, Vala, et al.
Pubblicazione: (2025)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
di: Behnia, Tina, et al.
Pubblicazione: (2025)
di: Behnia, Tina, et al.
Pubblicazione: (2025)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2026)
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2026)
On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy
di: Viswanathan, Karthik, et al.
Pubblicazione: (2025)
di: Viswanathan, Karthik, et al.
Pubblicazione: (2025)
Scale Determines Whether Language Models Organize Representation Geometry for Prediction
di: Xu, Weilun
Pubblicazione: (2026)
di: Xu, Weilun
Pubblicazione: (2026)
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
Cautious Next Token Prediction
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
di: Trauger, Jacob, et al.
Pubblicazione: (2025)
di: Trauger, Jacob, et al.
Pubblicazione: (2025)
Token Prediction as Implicit Classification to Identify LLM-Generated Text
di: Chen, Yutian, et al.
Pubblicazione: (2023)
di: Chen, Yutian, et al.
Pubblicazione: (2023)
Reasoning Bias of Next Token Prediction Training
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
ENTP: Encoder-only Next Token Prediction
di: Ewer, Ethan, et al.
Pubblicazione: (2024)
di: Ewer, Ethan, et al.
Pubblicazione: (2024)
Diversity or Precision? A Deep Dive into Next Token Prediction
di: Wu, Haoyuan, et al.
Pubblicazione: (2025)
di: Wu, Haoyuan, et al.
Pubblicazione: (2025)
How Language Directions Align with Token Geometry in Multilingual LLMs
di: Kim, JaeSeong, et al.
Pubblicazione: (2025)
di: Kim, JaeSeong, et al.
Pubblicazione: (2025)
SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
di: He, Yinhan, et al.
Pubblicazione: (2025)
di: He, Yinhan, et al.
Pubblicazione: (2025)
The Geometry of Tokens in Internal Representations of Large Language Models
di: Viswanathan, Karthik, et al.
Pubblicazione: (2025)
di: Viswanathan, Karthik, et al.
Pubblicazione: (2025)
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
di: Deng, Wenlong, et al.
Pubblicazione: (2024)
di: Deng, Wenlong, et al.
Pubblicazione: (2024)
Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension
di: Qi, Mengnan, et al.
Pubblicazione: (2024)
di: Qi, Mengnan, et al.
Pubblicazione: (2024)
Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semantics
di: R V, Kavin, et al.
Pubblicazione: (2025)
di: R V, Kavin, et al.
Pubblicazione: (2025)
Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction
di: Qian, Junlang, et al.
Pubblicazione: (2025)
di: Qian, Junlang, et al.
Pubblicazione: (2025)
Next Reply Prediction X Dataset: Linguistic Discrepancies in Naively Generated Content
di: Münker, Simon, et al.
Pubblicazione: (2026)
di: Münker, Simon, et al.
Pubblicazione: (2026)
YNTP-100: A Benchmark for Your Next Token Prediction with 100 People
di: Ding, Shiyao, et al.
Pubblicazione: (2025)
di: Ding, Shiyao, et al.
Pubblicazione: (2025)
Transformers as Support Vector Machines
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
di: Tarzanagh, Davoud Ataee, et al.
Pubblicazione: (2023)
Alternatives To Next Token Prediction In Text Generation -- A Survey
di: Wyatt, Charlie, et al.
Pubblicazione: (2025)
di: Wyatt, Charlie, et al.
Pubblicazione: (2025)
Fractal Patterns May Illuminate the Success of Next-Token Prediction
di: Alabdulmohsin, Ibrahim, et al.
Pubblicazione: (2024)
di: Alabdulmohsin, Ibrahim, et al.
Pubblicazione: (2024)
One Sentence, Two Embeddings: Contrastive Learning of Explicit and Implicit Semantic Representations
di: Oda, Kohei, et al.
Pubblicazione: (2025)
di: Oda, Kohei, et al.
Pubblicazione: (2025)
Modeling Next-Token Prediction as Left-Nested Intuitionistic Implication
di: Tarau, Paul
Pubblicazione: (2026)
di: Tarau, Paul
Pubblicazione: (2026)
How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them
di: Liao, Disen, et al.
Pubblicazione: (2026)
di: Liao, Disen, et al.
Pubblicazione: (2026)
How Muon's Spectral Design Benefits Generalization: A Study on Imbalanced Data
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
di: Vasudeva, Bhavya, et al.
Pubblicazione: (2025)
Mechanics of Next Token Prediction with Self-Attention
di: Li, Yingcong, et al.
Pubblicazione: (2024)
di: Li, Yingcong, et al.
Pubblicazione: (2024)
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
di: Schneider, Johannes
Pubblicazione: (2024)
di: Schneider, Johannes
Pubblicazione: (2024)
Linguistically Informed Tokenization Improves ASR for Underresourced Languages
di: Daul, Massimo, et al.
Pubblicazione: (2025)
di: Daul, Massimo, et al.
Pubblicazione: (2025)
Contextual Morphogenesis in Large Language Models: A Novel Approach to Self-Organizing Token Representations
di: Dombrowski, Alistair, et al.
Pubblicazione: (2025)
di: Dombrowski, Alistair, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
di: Zhao, Yize, et al.
Pubblicazione: (2024) -
Implicit Optimization Bias of Next-Token Prediction in Linear Models
di: Thrampoulidis, Christos
Pubblicazione: (2024) -
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
di: Deng, Wenlong, et al.
Pubblicazione: (2024) -
NITP: Next Implicit Token Prediction for LLM Pre-training
di: Zhang, Xiangdong, et al.
Pubblicazione: (2026) -
Why Loss Re-weighting Works If You Stop Early: Training Dynamics of Unconstrained Features
di: Zhao, Yize, et al.
Pubblicazione: (2026)