The pitfalls of next-token prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bachmann, Gregor, Nagarajan, Vaishnavh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
Language models are better than humans at next-token prediction
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
Deep sequence models tend to memorize geometrically; it is unclear why
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025)
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
Think before you speak: Training Language Models With Pause Tokens
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
Shaping capabilities with token-level data filtering
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
Scaling Transformer to 1M tokens and beyond with RMT
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
Interpretable Next-token Prediction via the Generalized Induction Head
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
Essential-Web v1.0: 24T tokens of organized web data
von: AI, Essential, et al.
Veröffentlicht: (2025)
von: AI, Essential, et al.
Veröffentlicht: (2025)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
von: Kumar, Anshul
Veröffentlicht: (2026)
von: Kumar, Anshul
Veröffentlicht: (2026)
On student-teacher deviations in distillation: does it pay to disobey?
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2023)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2023)
Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production
von: Galashov, Alexandre, et al.
Veröffentlicht: (2025)
von: Galashov, Alexandre, et al.
Veröffentlicht: (2025)
Task Facet Learning: A Structured Approach to Prompt Optimization
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024)
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
von: Witold, Waligóra
Veröffentlicht: (2024)
von: Witold, Waligóra
Veröffentlicht: (2024)
Surrogate modeling for interpreting black-box LLMs in medical predictions
von: Han, Changho, et al.
Veröffentlicht: (2026)
von: Han, Changho, et al.
Veröffentlicht: (2026)
NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
Transformers for molecular property prediction: Domain adaptation efficiently improves performance
von: Sultan, Afnan, et al.
Veröffentlicht: (2025)
von: Sultan, Afnan, et al.
Veröffentlicht: (2025)
On multi-token prediction for efficient LLM inference
von: Mehra, Somesh, et al.
Veröffentlicht: (2025)
von: Mehra, Somesh, et al.
Veröffentlicht: (2025)
Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language
von: Sharma, Yash, et al.
Veröffentlicht: (2024)
von: Sharma, Yash, et al.
Veröffentlicht: (2024)
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2024)
von: Pacchiardi, Lorenzo, et al.
Veröffentlicht: (2024)
Forma mentis networks predict creativity ratings of short texts via interpretable artificial intelligence in human and GPT-simulated raters
von: Haim, Edith, et al.
Veröffentlicht: (2024)
von: Haim, Edith, et al.
Veröffentlicht: (2024)
Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
Leveraging Large Language Models through Natural Language Processing to provide interpretable Machine Learning predictions of mental deterioration in real time
von: de Arriba-Pérez, Francisco, et al.
Veröffentlicht: (2024)
von: de Arriba-Pérez, Francisco, et al.
Veröffentlicht: (2024)
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
von: Burns, Thomas F, et al.
Veröffentlicht: (2025)
von: Burns, Thomas F, et al.
Veröffentlicht: (2025)
Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy
von: Penny-Dimri, Jahan C., et al.
Veröffentlicht: (2025)
von: Penny-Dimri, Jahan C., et al.
Veröffentlicht: (2025)
Understanding the Failure Modes of Out-of-Distribution Generalization
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2020)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2020)
Not all tokens are needed(NAT): token efficient reinforcement learning
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
von: Sang, Hejian, et al.
Veröffentlicht: (2026)
Large language models can accurately predict searcher preferences
von: Thomas, Paul, et al.
Veröffentlicht: (2023)
von: Thomas, Paul, et al.
Veröffentlicht: (2023)
A Family of LLMs Liberated from Static Vocabularies
von: Alpha, Aleph, et al.
Veröffentlicht: (2026)
von: Alpha, Aleph, et al.
Veröffentlicht: (2026)
Physics in Next-token Prediction
von: An, Hongjun, et al.
Veröffentlicht: (2024)
von: An, Hongjun, et al.
Veröffentlicht: (2024)
Where is the signal in tokenization space?
von: Geh, Renato Lui, et al.
Veröffentlicht: (2024)
von: Geh, Renato Lui, et al.
Veröffentlicht: (2024)
Command A: An Enterprise-Ready Large Language Model
von: Cohere, Team, et al.
Veröffentlicht: (2025)
von: Cohere, Team, et al.
Veröffentlicht: (2025)
Adaptive Pruning for Large Language Models with Structural Importance Awareness
von: Zheng, Haotian, et al.
Veröffentlicht: (2024)
von: Zheng, Haotian, et al.
Veröffentlicht: (2024)
Parameter Efficient Fine-tuning via Explained Variance Adaptation
von: Paischer, Fabian, et al.
Veröffentlicht: (2024)
von: Paischer, Fabian, et al.
Veröffentlicht: (2024)
When Can Transformers Count to n?
von: Yehudai, Gilad, et al.
Veröffentlicht: (2024)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2024)
PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency
von: Elements, Preferred, et al.
Veröffentlicht: (2024)
von: Elements, Preferred, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025) -
Language models are better than humans at next-token prediction
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022) -
Deep sequence models tend to memorize geometrically; it is unclear why
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025) -
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025) -
Think before you speak: Training Language Models With Pause Tokens
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)