Gespeichert in:
| Hauptverfasser: | Wu, Wilson, Morris, John X., Levine, Lionel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.00859 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visualizing token importance for black-box language models
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
von: Rauba, Paulius, et al.
Veröffentlicht: (2025)
Prompt reinforcing for long-term planning of large language models
von: Lin, Hsien-Chin, et al.
Veröffentlicht: (2025)
von: Lin, Hsien-Chin, et al.
Veröffentlicht: (2025)
EigenBench: A Comparative Behavioral Measure of Value Alignment
von: Chang, Jonathn, et al.
Veröffentlicht: (2025)
von: Chang, Jonathn, et al.
Veröffentlicht: (2025)
Where is the signal in tokenization space?
von: Geh, Renato Lui, et al.
Veröffentlicht: (2024)
von: Geh, Renato Lui, et al.
Veröffentlicht: (2024)
Extracting Prompts by Inverting LLM Outputs
von: Zhang, Collin, et al.
Veröffentlicht: (2024)
von: Zhang, Collin, et al.
Veröffentlicht: (2024)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
von: Kumar, Anshul
Veröffentlicht: (2026)
von: Kumar, Anshul
Veröffentlicht: (2026)
On multi-token prediction for efficient LLM inference
von: Mehra, Somesh, et al.
Veröffentlicht: (2025)
von: Mehra, Somesh, et al.
Veröffentlicht: (2025)
Language models are better than humans at next-token prediction
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
The pitfalls of next-token prediction
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
Do different prompting methods yield a common task representation in language models?
von: Davidson, Guy, et al.
Veröffentlicht: (2025)
von: Davidson, Guy, et al.
Veröffentlicht: (2025)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs
von: Merchant, Humzah, et al.
Veröffentlicht: (2025)
von: Merchant, Humzah, et al.
Veröffentlicht: (2025)
Byte-token Enhanced Language Models for Temporal Point Processes Analysis
von: Kong, Quyu, et al.
Veröffentlicht: (2025)
von: Kong, Quyu, et al.
Veröffentlicht: (2025)
Shaping capabilities with token-level data filtering
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
Scaling Transformer to 1M tokens and beyond with RMT
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
CogBench: a large language model walks into a psychology lab
von: Coda-Forno, Julian, et al.
Veröffentlicht: (2024)
von: Coda-Forno, Julian, et al.
Veröffentlicht: (2024)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
von: Zhao, Yize, et al.
Veröffentlicht: (2024)
von: Zhao, Yize, et al.
Veröffentlicht: (2024)
Large language models reorganize representational geometry during in-context learning
von: Xiong, Hua-Dong, et al.
Veröffentlicht: (2026)
von: Xiong, Hua-Dong, et al.
Veröffentlicht: (2026)
Interpretable Next-token Prediction via the Generalized Induction Head
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
Anchor function: a type of benchmark functions for studying language models
von: Zhang, Zhongwang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhongwang, et al.
Veröffentlicht: (2024)
Aligning language models with human preferences
von: Korbak, Tomasz
Veröffentlicht: (2024)
von: Korbak, Tomasz
Veröffentlicht: (2024)
Evaluating language models as risk scores
von: Cruz, André F., et al.
Veröffentlicht: (2024)
von: Cruz, André F., et al.
Veröffentlicht: (2024)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection
von: Shingi, Geet, et al.
Veröffentlicht: (2021)
von: Shingi, Geet, et al.
Veröffentlicht: (2021)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
Amortizing intractable inference in large language models
von: Hu, Edward J., et al.
Veröffentlicht: (2023)
von: Hu, Edward J., et al.
Veröffentlicht: (2023)
The language of time: a language model perspective on time-series foundation models
von: Xie, Yi, et al.
Veröffentlicht: (2025)
von: Xie, Yi, et al.
Veröffentlicht: (2025)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
Essential-Web v1.0: 24T tokens of organized web data
von: AI, Essential, et al.
Veröffentlicht: (2025)
von: AI, Essential, et al.
Veröffentlicht: (2025)
Perturbed examples reveal invariances shared by language models
von: Rawal, Ruchit, et al.
Veröffentlicht: (2023)
von: Rawal, Ruchit, et al.
Veröffentlicht: (2023)
A mean teacher algorithm for unlearning of language models
von: Klochkov, Yegor
Veröffentlicht: (2025)
von: Klochkov, Yegor
Veröffentlicht: (2025)
Learning to Detect Language Model Training Data via Active Reconstruction
von: Yin, Junjie Oscar, et al.
Veröffentlicht: (2026)
von: Yin, Junjie Oscar, et al.
Veröffentlicht: (2026)
Representation in large language models
von: Yetman, Cameron
Veröffentlicht: (2025)
von: Yetman, Cameron
Veröffentlicht: (2025)
Machine-generated text detection prevents language model collapse
von: Drayson, George, et al.
Veröffentlicht: (2025)
von: Drayson, George, et al.
Veröffentlicht: (2025)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
AI-AI Bias: large language models favor communications generated by large language models
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
von: Laurito, Walter, et al.
Veröffentlicht: (2024)
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
von: Raposo, David, et al.
Veröffentlicht: (2024)
von: Raposo, David, et al.
Veröffentlicht: (2024)
Simple linear attention language models balance the recall-throughput tradeoff
von: Arora, Simran, et al.
Veröffentlicht: (2024)
von: Arora, Simran, et al.
Veröffentlicht: (2024)
Just read twice: closing the recall gap for recurrent language models
von: Arora, Simran, et al.
Veröffentlicht: (2024)
von: Arora, Simran, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Visualizing token importance for black-box language models
von: Rauba, Paulius, et al.
Veröffentlicht: (2025) -
Prompt reinforcing for long-term planning of large language models
von: Lin, Hsien-Chin, et al.
Veröffentlicht: (2025) -
EigenBench: A Comparative Behavioral Measure of Value Alignment
von: Chang, Jonathn, et al.
Veröffentlicht: (2025) -
Where is the signal in tokenization space?
von: Geh, Renato Lui, et al.
Veröffentlicht: (2024) -
Extracting Prompts by Inverting LLM Outputs
von: Zhang, Collin, et al.
Veröffentlicht: (2024)