Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jia, Mumin, Diaz-Rodriguez, Jairo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mechanics of Next Token Prediction with Self-Attention
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
Consistent Kernel Change-Point Detection under m-Dependence for Text Segmentation
von: Diaz-Rodriguez, Jairo, et al.
Veröffentlicht: (2025)
von: Diaz-Rodriguez, Jairo, et al.
Veröffentlicht: (2025)
Unsupervised Text Segmentation via Kernel Change-Point Detection on Sentence Embeddings
von: Jia, Mumin, et al.
Veröffentlicht: (2026)
von: Jia, Mumin, et al.
Veröffentlicht: (2026)
When Can Digital Personas Reliably Approximate Human Survey Findings?
von: Jia, Mumin, et al.
Veröffentlicht: (2026)
von: Jia, Mumin, et al.
Veröffentlicht: (2026)
Cautious Next Token Prediction
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
A Law of Next-Token Prediction in Large Language Models
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
von: He, Hangfeng, et al.
Veröffentlicht: (2024)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
von: Shao, Chenze, et al.
Veröffentlicht: (2024)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
von: Rohekar, Raanan Y., et al.
Veröffentlicht: (2024)
von: Rohekar, Raanan Y., et al.
Veröffentlicht: (2024)
TokenButler: Token Importance is Predictable
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
Don't Shoot The Breeze: Topic Continuity Model Using Nonlinear Naive Bayes With Attention
von: Pi, Shu-Ting, et al.
Veröffentlicht: (2026)
von: Pi, Shu-Ting, et al.
Veröffentlicht: (2026)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
NextLocLLM: Location Semantics Modeling and Coordinate-Based Next Location Prediction with LLMs
von: Liu, Shuai, et al.
Veröffentlicht: (2024)
von: Liu, Shuai, et al.
Veröffentlicht: (2024)
TopicProphet: Prophesies on Temporal Topic Trends and Stocks
von: Kim, Olivia
Veröffentlicht: (2025)
von: Kim, Olivia
Veröffentlicht: (2025)
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
von: Julistiono, Addison Kristanto, et al.
Veröffentlicht: (2024)
von: Julistiono, Addison Kristanto, et al.
Veröffentlicht: (2024)
On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case Study
von: Alberghi, Riccardo, et al.
Veröffentlicht: (2025)
von: Alberghi, Riccardo, et al.
Veröffentlicht: (2025)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
von: Smith, Matthew L., et al.
Veröffentlicht: (2026)
von: Smith, Matthew L., et al.
Veröffentlicht: (2026)
Efficient Joint Prediction of Multiple Future Tokens
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2025)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2025)
Interpretable Next-token Prediction via the Generalized Induction Head
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
TopicDiff: A Topic-enriched Diffusion Approach for Multimodal Conversational Emotion Detection
von: Luo, Jiamin, et al.
Veröffentlicht: (2024)
von: Luo, Jiamin, et al.
Veröffentlicht: (2024)
Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document Generation
von: Yoo, YoungJoon, et al.
Veröffentlicht: (2023)
von: Yoo, YoungJoon, et al.
Veröffentlicht: (2023)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
Trainable Dynamic Mask Sparse Attention
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models
von: Kongmanee, Jaturong
Veröffentlicht: (2025)
von: Kongmanee, Jaturong
Veröffentlicht: (2025)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
TopicTag: Automatic Annotation of NMF Topic Models Using Chain of Thought and Prompt Tuning with LLMs
von: Wanna, Selma, et al.
Veröffentlicht: (2024)
von: Wanna, Selma, et al.
Veröffentlicht: (2024)
Empirical Capacity Model for Self-Attention Neural Networks
von: Härmä, Aki, et al.
Veröffentlicht: (2024)
von: Härmä, Aki, et al.
Veröffentlicht: (2024)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
Discovering Forbidden Topics in Language Models
von: Rager, Can, et al.
Veröffentlicht: (2025)
von: Rager, Can, et al.
Veröffentlicht: (2025)
Industry-Aligned Granular Topic Modeling
von: Moon, Sae Young, et al.
Veröffentlicht: (2026)
von: Moon, Sae Young, et al.
Veröffentlicht: (2026)
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
von: Guo, Zhenyuan, et al.
Veröffentlicht: (2026)
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
von: Kitouni, Ouail, et al.
Veröffentlicht: (2024)
von: Kitouni, Ouail, et al.
Veröffentlicht: (2024)
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
von: Zhong, Qimin, et al.
Veröffentlicht: (2026)
von: Zhong, Qimin, et al.
Veröffentlicht: (2026)
RecurFormer: Not All Transformer Heads Need Self-Attention
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
Counterfactual Token Generation in Large Language Models
von: Chatzi, Ivi, et al.
Veröffentlicht: (2024)
von: Chatzi, Ivi, et al.
Veröffentlicht: (2024)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
von: Kallini, Julie, et al.
Veröffentlicht: (2024)
Star Attention: Efficient LLM Inference over Long Sequences
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
von: Ding, Jiale, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mechanics of Next Token Prediction with Self-Attention
von: Li, Yingcong, et al.
Veröffentlicht: (2024) -
Consistent Kernel Change-Point Detection under m-Dependence for Text Segmentation
von: Diaz-Rodriguez, Jairo, et al.
Veröffentlicht: (2025) -
Unsupervised Text Segmentation via Kernel Change-Point Detection on Sentence Embeddings
von: Jia, Mumin, et al.
Veröffentlicht: (2026) -
When Can Digital Personas Reliably Approximate Human Survey Findings?
von: Jia, Mumin, et al.
Veröffentlicht: (2026) -
Cautious Next Token Prediction
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)