Prediction hubs are context-informed frequent tokens in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nielsen, Beatrix M. G., Macocco, Iuri, Baroni, Marco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Not a nuisance but a useful heuristic: Outlier dimensions favor frequent tokens in language models
von: Macocco, Iuri, et al.
Veröffentlicht: (2025)
von: Macocco, Iuri, et al.
Veröffentlicht: (2025)
Tracing Computation Density in LLMs
von: Kervadec, Corentin, et al.
Veröffentlicht: (2026)
von: Kervadec, Corentin, et al.
Veröffentlicht: (2026)
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
von: Gemini Team, et al.
Veröffentlicht: (2024)
von: Gemini Team, et al.
Veröffentlicht: (2024)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
von: Witold, Waligóra
Veröffentlicht: (2024)
von: Witold, Waligóra
Veröffentlicht: (2024)
Jacobian Scopes: token-level causal attributions in LLMs
von: Liu, Toni J. B., et al.
Veröffentlicht: (2026)
von: Liu, Toni J. B., et al.
Veröffentlicht: (2026)
Emergence of a High-Dimensional Abstraction Phase in Language Transformers
von: Cheng, Emily, et al.
Veröffentlicht: (2024)
von: Cheng, Emily, et al.
Veröffentlicht: (2024)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
von: Shrestha, Adarsha, et al.
Veröffentlicht: (2025)
von: Shrestha, Adarsha, et al.
Veröffentlicht: (2025)
Long-context LLMs Struggle with Long In-context Learning
von: Li, Tianle, et al.
Veröffentlicht: (2024)
von: Li, Tianle, et al.
Veröffentlicht: (2024)
MemoryPrompt: A Light Wrapper to Improve Context Tracking in Pre-trained Language Models
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
Unused information in token probability distribution of generative LLM: improving LLM reading comprehension through calculation of expected values
von: Zawistowski, Krystian
Veröffentlicht: (2024)
von: Zawistowski, Krystian
Veröffentlicht: (2024)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
The pitfalls of next-token prediction
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
ThoughtSource: A central hub for large language model reasoning data
von: Ott, Simon, et al.
Veröffentlicht: (2023)
von: Ott, Simon, et al.
Veröffentlicht: (2023)
On the token distance modeling ability of higher RoPE attention dimension
von: Hong, Xiangyu, et al.
Veröffentlicht: (2024)
von: Hong, Xiangyu, et al.
Veröffentlicht: (2024)
Scaled and Inter-token Relation Enhanced Transformer for Sample-restricted Residential NILM
von: Rahman, Minhajur, et al.
Veröffentlicht: (2024)
von: Rahman, Minhajur, et al.
Veröffentlicht: (2024)
Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
von: Hegde, Pradyoth, et al.
Veröffentlicht: (2025)
von: Hegde, Pradyoth, et al.
Veröffentlicht: (2025)
No Need for Explanations: LLMs can implicitly learn from mistakes in-context
von: Alazraki, Lisa, et al.
Veröffentlicht: (2025)
von: Alazraki, Lisa, et al.
Veröffentlicht: (2025)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
Shaping capabilities with token-level data filtering
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
DeepMLF: Multimodal language model with learnable tokens for deep fusion in sentiment analysis
von: Georgiou, Efthymios, et al.
Veröffentlicht: (2025)
von: Georgiou, Efthymios, et al.
Veröffentlicht: (2025)
SubTokenTest: A Practical Benchmark for Real-World Sub-token Understanding
von: Hou, Shuyang, et al.
Veröffentlicht: (2026)
von: Hou, Shuyang, et al.
Veröffentlicht: (2026)
Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives
von: Baker, Mohammed Abu, et al.
Veröffentlicht: (2026)
von: Baker, Mohammed Abu, et al.
Veröffentlicht: (2026)
Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
Scaling Transformer to 1M tokens and beyond with RMT
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
Evil twins are not that evil: Qualitative insights into machine-generated prompts
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
Do LLMs Dream of Ontologies?
von: Bombieri, Marco, et al.
Veröffentlicht: (2024)
von: Bombieri, Marco, et al.
Veröffentlicht: (2024)
Language models are better than humans at next-token prediction
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
von: Russell, Jenna, et al.
Veröffentlicht: (2025)
von: Russell, Jenna, et al.
Veröffentlicht: (2025)
A Decomposition Perspective to Long-context Reasoning for LLMs
von: Xiao, Yanling, et al.
Veröffentlicht: (2026)
von: Xiao, Yanling, et al.
Veröffentlicht: (2026)
Meaning Is Not A Metric: Using LLMs to make cultural context legible at scale
von: Kommers, Cody, et al.
Veröffentlicht: (2025)
von: Kommers, Cody, et al.
Veröffentlicht: (2025)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
TECP: Token-Entropy Conformal Prediction for LLMs
von: Xu, Beining, et al.
Veröffentlicht: (2025)
von: Xu, Beining, et al.
Veröffentlicht: (2025)
Redefining "Hallucination" in LLMs: Towards a psychology-informed framework for mitigating misinformation
von: Berberette, Elijah, et al.
Veröffentlicht: (2024)
von: Berberette, Elijah, et al.
Veröffentlicht: (2024)
Essential-Web v1.0: 24T tokens of organized web data
von: AI, Essential, et al.
Veröffentlicht: (2025)
von: AI, Essential, et al.
Veröffentlicht: (2025)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs
von: Chen, Ruxiao, et al.
Veröffentlicht: (2025)
von: Chen, Ruxiao, et al.
Veröffentlicht: (2025)
Harnessing LLMs for Educational Content-Driven Italian Crossword Generation
von: Zeinalipour, Kamyar, et al.
Veröffentlicht: (2024)
von: Zeinalipour, Kamyar, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Not a nuisance but a useful heuristic: Outlier dimensions favor frequent tokens in language models
von: Macocco, Iuri, et al.
Veröffentlicht: (2025) -
Tracing Computation Density in LLMs
von: Kervadec, Corentin, et al.
Veröffentlicht: (2026) -
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
von: Gemini Team, et al.
Veröffentlicht: (2024) -
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
von: Witold, Waligóra
Veröffentlicht: (2024) -
Jacobian Scopes: token-level causal attributions in LLMs
von: Liu, Toni J. B., et al.
Veröffentlicht: (2026)