Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points
Fuente:
arXiv
Saved in:
| Main Authors: | Varre, Aditya, Yüce, Gizem, Flammarion, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Parametric Distributions from Samples and Preferences
by: Jourdan, Marc, et al.
Published: (2025)
by: Jourdan, Marc, et al.
Published: (2025)
(How) Learning Rates Regulate Catastrophic Overtraining
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
by: Varre, Aditya, et al.
Published: (2026)
by: Varre, Aditya, et al.
Published: (2026)
Why Do We Need Weight Decay in Modern Deep Learning?
by: D'Angelo, Francesco, et al.
Published: (2023)
by: D'Angelo, Francesco, et al.
Published: (2023)
Transformers Can Represent $n$-gram Language Models
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
State space models can express n-gram languages
by: Nandakumar, Vinoth, et al.
Published: (2023)
by: Nandakumar, Vinoth, et al.
Published: (2023)
Understanding Transformers via N-gram Statistics
by: Nguyen, Timothy
Published: (2024)
by: Nguyen, Timothy
Published: (2024)
Faster Transformer Decoding: N-gram Masked Self-Attention
by: Chelba, Ciprian, et al.
Published: (2020)
by: Chelba, Ciprian, et al.
Published: (2020)
Transformers Learn Latent Mixture Models In-Context via Mirror Descent
by: D'Angelo, Francesco, et al.
Published: (2026)
by: D'Angelo, Francesco, et al.
Published: (2026)
Can Transformers Learn $n$-gram Language Models?
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Less is More: Understanding Word-level Textual Adversarial Attack via n-gram Frequency Descend
by: Lu, Ning, et al.
Published: (2023)
by: Lu, Ning, et al.
Published: (2023)
AdvSGM: Differentially Private Graph Learning via Adversarial Skip-gram Model
by: Zhang, Sen, et al.
Published: (2025)
by: Zhang, Sen, et al.
Published: (2025)
ECHOPulse: ECG controlled echocardio-grams video generation
by: Li, Yiwei, et al.
Published: (2024)
by: Li, Yiwei, et al.
Published: (2024)
Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
Incremental Learning of Sparse Attention Patterns in Transformers
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
Lossless Acceleration of Large Language Model via Adaptive N-gram Parallel Decoding
by: Ou, Jie, et al.
Published: (2024)
by: Ou, Jie, et al.
Published: (2024)
Early alignment in two-layer networks training is a two-edged sword
by: Boursier, Etienne, et al.
Published: (2024)
by: Boursier, Etienne, et al.
Published: (2024)
Simplicity bias and optimization threshold in two-layer ReLU networks
by: Boursier, Etienne, et al.
Published: (2024)
by: Boursier, Etienne, et al.
Published: (2024)
Penalising the biases in norm regularisation enforces sparsity
by: Boursier, Etienne, et al.
Published: (2023)
by: Boursier, Etienne, et al.
Published: (2023)
Selective Induction Heads: How Transformers Select Causal Structures In Context
by: D'Angelo, Francesco, et al.
Published: (2025)
by: D'Angelo, Francesco, et al.
Published: (2025)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
by: Boreiko, Valentyn, et al.
Published: (2024)
by: Boreiko, Valentyn, et al.
Published: (2024)
Learning Algorithms in the Limit
by: Papazov, Hristo, et al.
Published: (2025)
by: Papazov, Hristo, et al.
Published: (2025)
The Role of $n$-gram Smoothing in the Age of Neural Networks
by: Malagutti, Luca, et al.
Published: (2024)
by: Malagutti, Luca, et al.
Published: (2024)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Exact Learning of Arithmetic with Differentiable Agents
by: Papazov, Hristo, et al.
Published: (2025)
by: Papazov, Hristo, et al.
Published: (2025)
StateSMix: Online Lossless Compression via Mamba State Space Models and Sparse N-gram Context Mixing
by: Tacconelli, Roberto
Published: (2026)
by: Tacconelli, Roberto
Published: (2026)
HITgram: A Platform for Experimenting with n-gram Language Models
by: Dasgupta, Shibaranjani, et al.
Published: (2024)
by: Dasgupta, Shibaranjani, et al.
Published: (2024)
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
First-order ANIL provably learns representations despite overparametrization
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
by: Yüksel, Oğuz Kaan, et al.
Published: (2023)
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024)
by: Papazov, Hristo, et al.
Published: (2024)
Enhancing Bangla Language Next Word Prediction and Sentence Completion through Extended RNN with Bi-LSTM Model On N-gram Language
by: Islam, Md Robiul, et al.
Published: (2024)
by: Islam, Md Robiul, et al.
Published: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
Long-Context Linear System Identification
by: Yüksel, Oğuz Kaan, et al.
Published: (2024)
by: Yüksel, Oğuz Kaan, et al.
Published: (2024)
Beta-CoRM: A Bayesian Approach for $n$-gram Profiles Analysis
by: Perusquía, José A., et al.
Published: (2020)
by: Perusquía, José A., et al.
Published: (2020)
Limits of n-gram Style Control for LLMs via Logit-Space Injection
by: Ahmed, Sami-ul
Published: (2026)
by: Ahmed, Sami-ul
Published: (2026)
Evaluating n-gram Models for a Bilingual Word Sense Disambiguation Task
by: David Pinto
Published: (2011)
by: David Pinto
Published: (2011)
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
Similar Items
-
Learning Parametric Distributions from Samples and Preferences
by: Jourdan, Marc, et al.
Published: (2025) -
(How) Learning Rates Regulate Catastrophic Overtraining
by: Rofin, Mark, et al.
Published: (2026) -
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
by: Varre, Aditya, et al.
Published: (2026) -
Why Do We Need Weight Decay in Modern Deep Learning?
by: D'Angelo, Francesco, et al.
Published: (2023) -
Transformers Can Represent $n$-gram Language Models
by: Svete, Anej, et al.
Published: (2024)