On the scaling relationship between cloze probabilities and language model next-token prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Jacobs, Cassandra L., Grobol, Morgan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
di: Jacobs, Cassandra L., et al.
Pubblicazione: (2024)
di: Jacobs, Cassandra L., et al.
Pubblicazione: (2024)
A Bayesian account of pronoun and neopronoun acquisition
di: Jacobs, Cassandra L., et al.
Pubblicazione: (2025)
di: Jacobs, Cassandra L., et al.
Pubblicazione: (2025)
The pitfalls of next-token prediction
di: Bachmann, Gregor, et al.
Pubblicazione: (2024)
di: Bachmann, Gregor, et al.
Pubblicazione: (2024)
Language models are better than humans at next-token prediction
di: Shlegeris, Buck, et al.
Pubblicazione: (2022)
di: Shlegeris, Buck, et al.
Pubblicazione: (2022)
Less is more: Probabilistic reduction is best explained by small-scale predictability measures
di: Jacobs, Cassandra L., et al.
Pubblicazione: (2025)
di: Jacobs, Cassandra L., et al.
Pubblicazione: (2025)
Looking beyond the next token
di: Thankaraj, Abitha, et al.
Pubblicazione: (2025)
di: Thankaraj, Abitha, et al.
Pubblicazione: (2025)
Do language models plan ahead for future tokens?
di: Wu, Wilson, et al.
Pubblicazione: (2024)
di: Wu, Wilson, et al.
Pubblicazione: (2024)
Visualizing token importance for black-box language models
di: Rauba, Paulius, et al.
Pubblicazione: (2025)
di: Rauba, Paulius, et al.
Pubblicazione: (2025)
Subword Tokenization Strategies for Kurdish Word Embeddings
di: Salehi, Ali, et al.
Pubblicazione: (2025)
di: Salehi, Ali, et al.
Pubblicazione: (2025)
Collaborative decoding of critical tokens for boosting factuality of large language models
di: Jin, Lifeng, et al.
Pubblicazione: (2024)
di: Jin, Lifeng, et al.
Pubblicazione: (2024)
The distribution of discourse relations within and across turns in spontaneous conversation
di: Cortez, S. Magalí López, et al.
Pubblicazione: (2023)
di: Cortez, S. Magalí López, et al.
Pubblicazione: (2023)
Comparative analysis of subword tokenization approaches for Indian languages
di: Das, Sudhansu Bala, et al.
Pubblicazione: (2025)
di: Das, Sudhansu Bala, et al.
Pubblicazione: (2025)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
di: Nagarajan, Vaishnavh, et al.
Pubblicazione: (2025)
di: Nagarajan, Vaishnavh, et al.
Pubblicazione: (2025)
Revisiting subword tokenization: A case study on affixal negation in large language models
di: Truong, Thinh Hung, et al.
Pubblicazione: (2024)
di: Truong, Thinh Hung, et al.
Pubblicazione: (2024)
DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis
di: Georgiou, Efthymios, et al.
Pubblicazione: (2025)
di: Georgiou, Efthymios, et al.
Pubblicazione: (2025)
On multi-token prediction for efficient LLM inference
di: Mehra, Somesh, et al.
Pubblicazione: (2025)
di: Mehra, Somesh, et al.
Pubblicazione: (2025)
Large Language Model probabilities cannot distinguish between possible and impossible language
di: Leivada, Evelina, et al.
Pubblicazione: (2025)
di: Leivada, Evelina, et al.
Pubblicazione: (2025)
Finetuning LLMs for EvaCun 2025 token prediction shared task
di: Jon, Josef, et al.
Pubblicazione: (2025)
di: Jon, Josef, et al.
Pubblicazione: (2025)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
di: Witold, Waligóra
Pubblicazione: (2024)
di: Witold, Waligóra
Pubblicazione: (2024)
Logical forms complement probability in understanding language model (and human) performance
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
Contextual morphologically-guided tokenization for Latin encoder models
di: Hudspeth, Marisa, et al.
Pubblicazione: (2025)
di: Hudspeth, Marisa, et al.
Pubblicazione: (2025)
Not a nuisance but a useful heuristic: Outlier dimensions favor frequent tokens in language models
di: Macocco, Iuri, et al.
Pubblicazione: (2025)
di: Macocco, Iuri, et al.
Pubblicazione: (2025)
DataComp-LM: In search of the next generation of training sets for language models
di: Li, Jeffrey, et al.
Pubblicazione: (2024)
di: Li, Jeffrey, et al.
Pubblicazione: (2024)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
di: Shrestha, Adarsha, et al.
Pubblicazione: (2025)
di: Shrestha, Adarsha, et al.
Pubblicazione: (2025)
Multilingual acoustic word embeddings for zero-resource languages
di: Jacobs, Christiaan
Pubblicazione: (2024)
di: Jacobs, Christiaan
Pubblicazione: (2024)
Unused information in token probability distribution of generative LLM: improving LLM reading comprehension through calculation of expected values
di: Zawistowski, Krystian
Pubblicazione: (2024)
di: Zawistowski, Krystian
Pubblicazione: (2024)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
di: Gubian, Michele, et al.
Pubblicazione: (2025)
di: Gubian, Michele, et al.
Pubblicazione: (2025)
Complete asymptotic type-token relationship for growing complex systems with inverse power-law count rankings
di: Rosillo-Rodes, Pablo, et al.
Pubblicazione: (2025)
di: Rosillo-Rodes, Pablo, et al.
Pubblicazione: (2025)
Meta predictive learning model of languages in neural circuits
di: Li, Chan, et al.
Pubblicazione: (2023)
di: Li, Chan, et al.
Pubblicazione: (2023)
Practical token pruning for foundation models in few-shot conversational virtual assistant systems
di: Qi, Haode, et al.
Pubblicazione: (2024)
di: Qi, Haode, et al.
Pubblicazione: (2024)
Why do LLMs attend to the first token?
di: Barbero, Federico, et al.
Pubblicazione: (2025)
di: Barbero, Federico, et al.
Pubblicazione: (2025)
Where is the signal in tokenization space?
di: Geh, Renato Lui, et al.
Pubblicazione: (2024)
di: Geh, Renato Lui, et al.
Pubblicazione: (2024)
On the token distance modeling ability of higher RoPE attention dimension
di: Hong, Xiangyu, et al.
Pubblicazione: (2024)
di: Hong, Xiangyu, et al.
Pubblicazione: (2024)
Large-scale moral machine experiment on large language models
di: Ahmad, Muhammad Shahrul Zaim bin, et al.
Pubblicazione: (2024)
di: Ahmad, Muhammad Shahrul Zaim bin, et al.
Pubblicazione: (2024)
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
di: Yang, Tianle, et al.
Pubblicazione: (2026)
di: Yang, Tianle, et al.
Pubblicazione: (2026)
Is my model "mind blurting"? Interpreting the dynamics of reasoning tokens with Recurrence Quantification Analysis (RQA)
di: Pham, Quoc Tuan, et al.
Pubblicazione: (2026)
di: Pham, Quoc Tuan, et al.
Pubblicazione: (2026)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
di: Kumar, Anshul
Pubblicazione: (2026)
di: Kumar, Anshul
Pubblicazione: (2026)
Interpreting token compositionality in LLMs: A robustness analysis
di: Aljaafari, Nura, et al.
Pubblicazione: (2024)
di: Aljaafari, Nura, et al.
Pubblicazione: (2024)
Language models and brains align due to more than next-word prediction and word-level information
di: Merlin, Gabriele, et al.
Pubblicazione: (2022)
di: Merlin, Gabriele, et al.
Pubblicazione: (2022)
When can isotropy help adapt LLMs' next word prediction to numerical domains?
di: Shelim, Rashed, et al.
Pubblicazione: (2025)
di: Shelim, Rashed, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
di: Jacobs, Cassandra L., et al.
Pubblicazione: (2024) -
A Bayesian account of pronoun and neopronoun acquisition
di: Jacobs, Cassandra L., et al.
Pubblicazione: (2025) -
The pitfalls of next-token prediction
di: Bachmann, Gregor, et al.
Pubblicazione: (2024) -
Language models are better than humans at next-token prediction
di: Shlegeris, Buck, et al.
Pubblicazione: (2022) -
Less is more: Probabilistic reduction is best explained by small-scale predictability measures
di: Jacobs, Cassandra L., et al.
Pubblicazione: (2025)