Saved in:
| Main Authors: | Jacobs, Cassandra L., Grobol, Morgan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.17848 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
by: Jacobs, Cassandra L., et al.
Published: (2024)
by: Jacobs, Cassandra L., et al.
Published: (2024)
A Bayesian account of pronoun and neopronoun acquisition
by: Jacobs, Cassandra L., et al.
Published: (2025)
by: Jacobs, Cassandra L., et al.
Published: (2025)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
Less is more: Probabilistic reduction is best explained by small-scale predictability measures
by: Jacobs, Cassandra L., et al.
Published: (2025)
by: Jacobs, Cassandra L., et al.
Published: (2025)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
Subword Tokenization Strategies for Kurdish Word Embeddings
by: Salehi, Ali, et al.
Published: (2025)
by: Salehi, Ali, et al.
Published: (2025)
Do language models plan ahead for future tokens?
by: Wu, Wilson, et al.
Published: (2024)
by: Wu, Wilson, et al.
Published: (2024)
Visualizing token importance for black-box language models
by: Rauba, Paulius, et al.
Published: (2025)
by: Rauba, Paulius, et al.
Published: (2025)
Collaborative decoding of critical tokens for boosting factuality of large language models
by: Jin, Lifeng, et al.
Published: (2024)
by: Jin, Lifeng, et al.
Published: (2024)
The distribution of discourse relations within and across turns in spontaneous conversation
by: Cortez, S. Magalí López, et al.
Published: (2023)
by: Cortez, S. Magalí López, et al.
Published: (2023)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
Comparative analysis of subword tokenization approaches for Indian languages
by: Das, Sudhansu Bala, et al.
Published: (2025)
by: Das, Sudhansu Bala, et al.
Published: (2025)
Revisiting subword tokenization: A case study on affixal negation in large language models
by: Truong, Thinh Hung, et al.
Published: (2024)
by: Truong, Thinh Hung, et al.
Published: (2024)
DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis
by: Georgiou, Efthymios, et al.
Published: (2025)
by: Georgiou, Efthymios, et al.
Published: (2025)
On multi-token prediction for efficient LLM inference
by: Mehra, Somesh, et al.
Published: (2025)
by: Mehra, Somesh, et al.
Published: (2025)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
by: Witold, Waligóra
Published: (2024)
by: Witold, Waligóra
Published: (2024)
Large Language Model probabilities cannot distinguish between possible and impossible language
by: Leivada, Evelina, et al.
Published: (2025)
by: Leivada, Evelina, et al.
Published: (2025)
Finetuning LLMs for EvaCun 2025 token prediction shared task
by: Jon, Josef, et al.
Published: (2025)
by: Jon, Josef, et al.
Published: (2025)
Logical forms complement probability in understanding language model (and human) performance
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Not a nuisance but a useful heuristic: Outlier dimensions favor frequent tokens in language models
by: Macocco, Iuri, et al.
Published: (2025)
by: Macocco, Iuri, et al.
Published: (2025)
Contextual morphologically-guided tokenization for Latin encoder models
by: Hudspeth, Marisa, et al.
Published: (2025)
by: Hudspeth, Marisa, et al.
Published: (2025)
DataComp-LM: In search of the next generation of training sets for language models
by: Li, Jeffrey, et al.
Published: (2024)
by: Li, Jeffrey, et al.
Published: (2024)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
Multilingual acoustic word embeddings for zero-resource languages
by: Jacobs, Christiaan
Published: (2024)
by: Jacobs, Christiaan
Published: (2024)
Unused information in token probability distribution of generative LLM: improving LLM reading comprehension through calculation of expected values
by: Zawistowski, Krystian
Published: (2024)
by: Zawistowski, Krystian
Published: (2024)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
by: Gubian, Michele, et al.
Published: (2025)
by: Gubian, Michele, et al.
Published: (2025)
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
by: Yang, Tianle, et al.
Published: (2026)
by: Yang, Tianle, et al.
Published: (2026)
Complete asymptotic type-token relationship for growing complex systems with inverse power-law count rankings
by: Rosillo-Rodes, Pablo, et al.
Published: (2025)
by: Rosillo-Rodes, Pablo, et al.
Published: (2025)
Meta predictive learning model of languages in neural circuits
by: Li, Chan, et al.
Published: (2023)
by: Li, Chan, et al.
Published: (2023)
Where is the signal in tokenization space?
by: Geh, Renato Lui, et al.
Published: (2024)
by: Geh, Renato Lui, et al.
Published: (2024)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
by: Kumar, Anshul
Published: (2026)
by: Kumar, Anshul
Published: (2026)
On the token distance modeling ability of higher RoPE attention dimension
by: Hong, Xiangyu, et al.
Published: (2024)
by: Hong, Xiangyu, et al.
Published: (2024)
Large-scale moral machine experiment on large language models
by: Ahmad, Muhammad Shahrul Zaim bin, et al.
Published: (2024)
by: Ahmad, Muhammad Shahrul Zaim bin, et al.
Published: (2024)
Why do LLMs attend to the first token?
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Practical token pruning for foundation models in few-shot conversational virtual assistant systems
by: Qi, Haode, et al.
Published: (2024)
by: Qi, Haode, et al.
Published: (2024)
Antisocial behavior towards large language model users: experimental evidence
by: Niszczota, Paweł, et al.
Published: (2026)
by: Niszczota, Paweł, et al.
Published: (2026)
Analyzing Large language models chatbots: An experimental approach using a probability test
by: Peruchini, Melise, et al.
Published: (2024)
by: Peruchini, Melise, et al.
Published: (2024)
Is my model "mind blurting"? Interpreting the dynamics of reasoning tokens with Recurrence Quantification Analysis (RQA)
by: Pham, Quoc Tuan, et al.
Published: (2026)
by: Pham, Quoc Tuan, et al.
Published: (2026)
Language models and brains align due to more than next-word prediction and word-level information
by: Merlin, Gabriele, et al.
Published: (2022)
by: Merlin, Gabriele, et al.
Published: (2022)
Similar Items
-
Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
by: Jacobs, Cassandra L., et al.
Published: (2024) -
A Bayesian account of pronoun and neopronoun acquisition
by: Jacobs, Cassandra L., et al.
Published: (2025) -
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024) -
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022) -
Less is more: Probabilistic reduction is best explained by small-scale predictability measures
by: Jacobs, Cassandra L., et al.
Published: (2025)