Language models are better than humans at next-token prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shlegeris, Buck, Roger, Fabien, Chan, Lawrence, McLean, Euan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The pitfalls of next-token prediction
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
Exploiting Novel GPT-4 APIs
von: Pelrine, Kellin, et al.
Veröffentlicht: (2023)
von: Pelrine, Kellin, et al.
Veröffentlicht: (2023)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
von: Guo, Shiyuan, et al.
Veröffentlicht: (2025)
von: Guo, Shiyuan, et al.
Veröffentlicht: (2025)
Alignment faking in large language models
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
von: Greenblatt, Ryan, et al.
Veröffentlicht: (2024)
Language models show human-like content effects on reasoning tasks
von: Dasgupta, Ishita, et al.
Veröffentlicht: (2022)
von: Dasgupta, Ishita, et al.
Veröffentlicht: (2022)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
von: Nunez, Jeanmely Rojas, et al.
Veröffentlicht: (2026)
von: Nunez, Jeanmely Rojas, et al.
Veröffentlicht: (2026)
Two are better than one: Context window extension with multi-grained self-injection
von: Han, Wei, et al.
Veröffentlicht: (2024)
von: Han, Wei, et al.
Veröffentlicht: (2024)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics
von: Alakuijala, Minttu, et al.
Veröffentlicht: (2024)
von: Alakuijala, Minttu, et al.
Veröffentlicht: (2024)
Shaping capabilities with token-level data filtering
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
Can Go AIs be adversarially robust?
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
von: Tseng, Tom, et al.
Veröffentlicht: (2024)
Scaling Transformer to 1M tokens and beyond with RMT
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
Interpretable Next-token Prediction via the Generalized Induction Head
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
A NotSo Simple Way to Beat Simple Bench
von: Sane, Soham, et al.
Veröffentlicht: (2024)
von: Sane, Soham, et al.
Veröffentlicht: (2024)
Self-Improving Pretraining: using post-trained models to pretrain better models
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
Essential-Web v1.0: 24T tokens of organized web data
von: AI, Essential, et al.
Veröffentlicht: (2025)
von: AI, Essential, et al.
Veröffentlicht: (2025)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
TPTT: Transforming Pretrained Transformers into Titans
von: Furfaro, Fabien
Veröffentlicht: (2025)
von: Furfaro, Fabien
Veröffentlicht: (2025)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
von: Mallen, Alex, et al.
Veröffentlicht: (2024)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
von: Griffin, Charlie, et al.
Veröffentlicht: (2024)
von: Griffin, Charlie, et al.
Veröffentlicht: (2024)
Polysemanticity and Capacity in Neural Networks
von: Scherlis, Adam, et al.
Veröffentlicht: (2022)
von: Scherlis, Adam, et al.
Veröffentlicht: (2022)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
von: Kumar, Anshul
Veröffentlicht: (2026)
von: Kumar, Anshul
Veröffentlicht: (2026)
Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models
von: Yan, Xue, et al.
Veröffentlicht: (2023)
von: Yan, Xue, et al.
Veröffentlicht: (2023)
Is continuous CoT better suited for multi-lingual reasoning?
von: Bashir, Ali Hamza, et al.
Veröffentlicht: (2026)
von: Bashir, Ali Hamza, et al.
Veröffentlicht: (2026)
Post-training makes large language models less human-like
von: Binz, Marcel, et al.
Veröffentlicht: (2026)
von: Binz, Marcel, et al.
Veröffentlicht: (2026)
Auditing language models for hidden objectives
von: Marks, Samuel, et al.
Veröffentlicht: (2025)
von: Marks, Samuel, et al.
Veröffentlicht: (2025)
Forma mentis networks predict creativity ratings of short texts via interpretable artificial intelligence in human and GPT-simulated raters
von: Haim, Edith, et al.
Veröffentlicht: (2024)
von: Haim, Edith, et al.
Veröffentlicht: (2024)
Compositional Steering of Large Language Models with Steering Tokens
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
von: Radevski, Gorjan, et al.
Veröffentlicht: (2026)
More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2025)
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Surrogate modeling for interpreting black-box LLMs in medical predictions
von: Han, Changho, et al.
Veröffentlicht: (2026)
von: Han, Changho, et al.
Veröffentlicht: (2026)
GAPMAP: Mapping Scientific Knowledge Gaps in Biomedical Literature Using Large Language Models
von: Salem, Nourah M, et al.
Veröffentlicht: (2025)
von: Salem, Nourah M, et al.
Veröffentlicht: (2025)
Evaluating Language-Model Agents on Realistic Autonomous Tasks
von: Kinniment, Megan, et al.
Veröffentlicht: (2023)
von: Kinniment, Megan, et al.
Veröffentlicht: (2023)
Large Language Models Assume People are More Rational than We Really are
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
Enriching language models with graph-based context information to better understand textual data
von: Roethel, Albert, et al.
Veröffentlicht: (2023)
von: Roethel, Albert, et al.
Veröffentlicht: (2023)
Training Language Models with Language Feedback at Scale
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo
von: Zhao, Stephen, et al.
Veröffentlicht: (2024)
von: Zhao, Stephen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The pitfalls of next-token prediction
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024) -
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025) -
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025) -
Exploiting Novel GPT-4 APIs
von: Pelrine, Kellin, et al.
Veröffentlicht: (2023) -
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
von: Guo, Shiyuan, et al.
Veröffentlicht: (2025)