Do language models plan ahead for future tokens?
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Wilson, Morris, John X., Levine, Lionel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Visualizing token importance for black-box language models
di: Rauba, Paulius, et al.
Pubblicazione: (2025)
di: Rauba, Paulius, et al.
Pubblicazione: (2025)
Prompt reinforcing for long-term planning of large language models
di: Lin, Hsien-Chin, et al.
Pubblicazione: (2025)
di: Lin, Hsien-Chin, et al.
Pubblicazione: (2025)
Where is the signal in tokenization space?
di: Geh, Renato Lui, et al.
Pubblicazione: (2024)
di: Geh, Renato Lui, et al.
Pubblicazione: (2024)
EigenBench: A Comparative Behavioral Measure of Value Alignment
di: Chang, Jonathn, et al.
Pubblicazione: (2025)
di: Chang, Jonathn, et al.
Pubblicazione: (2025)
Extracting Prompts by Inverting LLM Outputs
di: Zhang, Collin, et al.
Pubblicazione: (2024)
di: Zhang, Collin, et al.
Pubblicazione: (2024)
On multi-token prediction for efficient LLM inference
di: Mehra, Somesh, et al.
Pubblicazione: (2025)
di: Mehra, Somesh, et al.
Pubblicazione: (2025)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
di: Kumar, Anshul
Pubblicazione: (2026)
di: Kumar, Anshul
Pubblicazione: (2026)
Do different prompting methods yield a common task representation in language models?
di: Davidson, Guy, et al.
Pubblicazione: (2025)
di: Davidson, Guy, et al.
Pubblicazione: (2025)
Language models are better than humans at next-token prediction
di: Shlegeris, Buck, et al.
Pubblicazione: (2022)
di: Shlegeris, Buck, et al.
Pubblicazione: (2022)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
di: Singh, Aaditya K., et al.
Pubblicazione: (2024)
The pitfalls of next-token prediction
di: Bachmann, Gregor, et al.
Pubblicazione: (2024)
di: Bachmann, Gregor, et al.
Pubblicazione: (2024)
Looking beyond the next token
di: Thankaraj, Abitha, et al.
Pubblicazione: (2025)
di: Thankaraj, Abitha, et al.
Pubblicazione: (2025)
Byte-token Enhanced Language Models for Temporal Point Processes Analysis
di: Kong, Quyu, et al.
Pubblicazione: (2025)
di: Kong, Quyu, et al.
Pubblicazione: (2025)
A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs
di: Merchant, Humzah, et al.
Pubblicazione: (2025)
di: Merchant, Humzah, et al.
Pubblicazione: (2025)
Shaping capabilities with token-level data filtering
di: Rathi, Neil, et al.
Pubblicazione: (2026)
di: Rathi, Neil, et al.
Pubblicazione: (2026)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
di: Zhao, Yize, et al.
Pubblicazione: (2024)
di: Zhao, Yize, et al.
Pubblicazione: (2024)
Scaling Transformer to 1M tokens and beyond with RMT
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
di: Bulatov, Aydar, et al.
Pubblicazione: (2023)
Anchor function: a type of benchmark functions for studying language models
di: Zhang, Zhongwang, et al.
Pubblicazione: (2024)
di: Zhang, Zhongwang, et al.
Pubblicazione: (2024)
Aligning language models with human preferences
di: Korbak, Tomasz
Pubblicazione: (2024)
di: Korbak, Tomasz
Pubblicazione: (2024)
Evaluating language models as risk scores
di: Cruz, André F., et al.
Pubblicazione: (2024)
di: Cruz, André F., et al.
Pubblicazione: (2024)
CogBench: a large language model walks into a psychology lab
di: Coda-Forno, Julian, et al.
Pubblicazione: (2024)
di: Coda-Forno, Julian, et al.
Pubblicazione: (2024)
AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection
di: Shingi, Geet, et al.
Pubblicazione: (2021)
di: Shingi, Geet, et al.
Pubblicazione: (2021)
Large language models reorganize representational geometry during in-context learning
di: Xiong, Hua-Dong, et al.
Pubblicazione: (2026)
di: Xiong, Hua-Dong, et al.
Pubblicazione: (2026)
Amortizing intractable inference in large language models
di: Hu, Edward J., et al.
Pubblicazione: (2023)
di: Hu, Edward J., et al.
Pubblicazione: (2023)
Interpretable Next-token Prediction via the Generalized Induction Head
di: Kim, Eunji, et al.
Pubblicazione: (2024)
di: Kim, Eunji, et al.
Pubblicazione: (2024)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
di: Nagarajan, Vaishnavh, et al.
Pubblicazione: (2025)
di: Nagarajan, Vaishnavh, et al.
Pubblicazione: (2025)
Perturbed examples reveal invariances shared by language models
di: Rawal, Ruchit, et al.
Pubblicazione: (2023)
di: Rawal, Ruchit, et al.
Pubblicazione: (2023)
A mean teacher algorithm for unlearning of language models
di: Klochkov, Yegor
Pubblicazione: (2025)
di: Klochkov, Yegor
Pubblicazione: (2025)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
di: Kwek, Eugene, et al.
Pubblicazione: (2025)
di: Kwek, Eugene, et al.
Pubblicazione: (2025)
The language of time: a language model perspective on time-series foundation models
di: Xie, Yi, et al.
Pubblicazione: (2025)
di: Xie, Yi, et al.
Pubblicazione: (2025)
Machine-generated text detection prevents language model collapse
di: Drayson, George, et al.
Pubblicazione: (2025)
di: Drayson, George, et al.
Pubblicazione: (2025)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
di: Marconato, Emanuele, et al.
Pubblicazione: (2024)
di: Marconato, Emanuele, et al.
Pubblicazione: (2024)
Essential-Web v1.0: 24T tokens of organized web data
di: AI, Essential, et al.
Pubblicazione: (2025)
di: AI, Essential, et al.
Pubblicazione: (2025)
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
di: Raposo, David, et al.
Pubblicazione: (2024)
di: Raposo, David, et al.
Pubblicazione: (2024)
Simple linear attention language models balance the recall-throughput tradeoff
di: Arora, Simran, et al.
Pubblicazione: (2024)
di: Arora, Simran, et al.
Pubblicazione: (2024)
Just read twice: closing the recall gap for recurrent language models
di: Arora, Simran, et al.
Pubblicazione: (2024)
di: Arora, Simran, et al.
Pubblicazione: (2024)
Zero-shot generation of synthetic neurosurgical data with large language models
di: Barr, Austin A., et al.
Pubblicazione: (2025)
di: Barr, Austin A., et al.
Pubblicazione: (2025)
Repetitions are not all alike: distinct mechanisms sustain repetition in language models
di: Mahaut, Matéo, et al.
Pubblicazione: (2025)
di: Mahaut, Matéo, et al.
Pubblicazione: (2025)
Only relative ranks matter in weight-clustered large language models
di: Aizpurua, Borja, et al.
Pubblicazione: (2026)
di: Aizpurua, Borja, et al.
Pubblicazione: (2026)
How do language models learn facts? Dynamics, curricula and hallucinations
di: Zucchet, Nicolas, et al.
Pubblicazione: (2025)
di: Zucchet, Nicolas, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Visualizing token importance for black-box language models
di: Rauba, Paulius, et al.
Pubblicazione: (2025) -
Prompt reinforcing for long-term planning of large language models
di: Lin, Hsien-Chin, et al.
Pubblicazione: (2025) -
Where is the signal in tokenization space?
di: Geh, Renato Lui, et al.
Pubblicazione: (2024) -
EigenBench: A Comparative Behavioral Measure of Value Alignment
di: Chang, Jonathn, et al.
Pubblicazione: (2025) -
Extracting Prompts by Inverting LLM Outputs
di: Zhang, Collin, et al.
Pubblicazione: (2024)