All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Marconato, Emanuele, Lachapelle, Sébastien, Weichwald, Sebastian, Gresele, Luigi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
by: Joshi, Shruti, et al.
Published: (2025)
by: Joshi, Shruti, et al.
Published: (2025)
When Does Closeness in Distribution Imply Representational Similarity? An Identifiability Perspective
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
What is causal about causal models and representations?
by: Jørgensen, Frederik Hytting, et al.
Published: (2025)
by: Jørgensen, Frederik Hytting, et al.
Published: (2025)
Relational Linear Properties in Language Models: An Empirical Investigation
by: Valer, Giovanni, et al.
Published: (2026)
by: Valer, Giovanni, et al.
Published: (2026)
Logit Distance Bounds Representational Similarity
by: Nielsen, Beatrix M. G., et al.
Published: (2026)
by: Nielsen, Beatrix M. G., et al.
Published: (2026)
Interpretable Next-token Prediction via the Generalized Induction Head
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
CLadder: Assessing Causal Reasoning in Language Models
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Identifiability of Potentially Degenerate Gaussian Mixture Models With Piecewise Affine Mixing
by: Xu, Danru, et al.
Published: (2026)
by: Xu, Danru, et al.
Published: (2026)
Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
by: Braun, Joschka
Published: (2026)
by: Braun, Joschka
Published: (2026)
Language models are better than humans at next-token prediction
by: Shlegeris, Buck, et al.
Published: (2022)
by: Shlegeris, Buck, et al.
Published: (2022)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
by: Kwek, Eugene, et al.
Published: (2025)
by: Kwek, Eugene, et al.
Published: (2025)
Shaping capabilities with token-level data filtering
by: Rathi, Neil, et al.
Published: (2026)
by: Rathi, Neil, et al.
Published: (2026)
Scaling Transformer to 1M tokens and beyond with RMT
by: Bulatov, Aydar, et al.
Published: (2023)
by: Bulatov, Aydar, et al.
Published: (2023)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
All Language Models Large and Small
by: Chen, Zhixun, et al.
Published: (2024)
by: Chen, Zhixun, et al.
Published: (2024)
On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case Study
by: Alberghi, Riccardo, et al.
Published: (2025)
by: Alberghi, Riccardo, et al.
Published: (2025)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
A Law of Next-Token Prediction in Large Language Models
by: He, Hangfeng, et al.
Published: (2024)
by: He, Hangfeng, et al.
Published: (2024)
Shortcuts and Identifiability in Concept-based Models from a Neuro-Symbolic Lens
by: Bortolotti, Samuele, et al.
Published: (2025)
by: Bortolotti, Samuele, et al.
Published: (2025)
Essential-Web v1.0: 24T tokens of organized web data
by: AI, Essential, et al.
Published: (2025)
by: AI, Essential, et al.
Published: (2025)
One Model for All: Multi-Objective Controllable Language Models
by: He, Qiang, et al.
Published: (2026)
by: He, Qiang, et al.
Published: (2026)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
Pruning Large Language Models by Identifying and Preserving Functional Networks
by: Liu, Yiheng, et al.
Published: (2025)
by: Liu, Yiheng, et al.
Published: (2025)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
by: Shao, Chenze, et al.
Published: (2024)
by: Shao, Chenze, et al.
Published: (2024)
Equivalent Linear Mappings of Large Language Models
by: Golden, James R.
Published: (2025)
by: Golden, James R.
Published: (2025)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
by: Guo, Shiyuan, et al.
Published: (2025)
by: Guo, Shiyuan, et al.
Published: (2025)
Preserving Privacy in Large Language Models: A Survey on Current Threats and Solutions
by: Miranda, Michele, et al.
Published: (2024)
by: Miranda, Michele, et al.
Published: (2024)
Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models
by: Kongmanee, Jaturong
Published: (2025)
by: Kongmanee, Jaturong
Published: (2025)
Communicating Sound Through Natural Language
by: Rossi, Emanuele, et al.
Published: (2026)
by: Rossi, Emanuele, et al.
Published: (2026)
Parallax: Parameterized Local Linear Attention for Language Modeling
by: Zuo, Yifei, et al.
Published: (2026)
by: Zuo, Yifei, et al.
Published: (2026)
The Linear Representation Hypothesis and the Geometry of Large Language Models
by: Park, Kiho, et al.
Published: (2023)
by: Park, Kiho, et al.
Published: (2023)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
by: Kumar, Anshul
Published: (2026)
by: Kumar, Anshul
Published: (2026)
Small Language Models for Application Interactions: A Case Study
by: Li, Beibin, et al.
Published: (2024)
by: Li, Beibin, et al.
Published: (2024)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
On the Identifiability of Latent Action Policies
by: Lachapelle, Sébastien
Published: (2025)
by: Lachapelle, Sébastien
Published: (2025)
Train Once, Answer All: Many Pretraining Experiments for the Cost of One
by: Bordt, Sebastian, et al.
Published: (2025)
by: Bordt, Sebastian, et al.
Published: (2025)
NextLocLLM: Location Semantics Modeling and Coordinate-Based Next Location Prediction with LLMs
by: Liu, Shuai, et al.
Published: (2024)
by: Liu, Shuai, et al.
Published: (2024)
Similar Items
-
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
by: Joshi, Shruti, et al.
Published: (2025) -
When Does Closeness in Distribution Imply Representational Similarity? An Identifiability Perspective
by: Nielsen, Beatrix M. G., et al.
Published: (2025) -
What is causal about causal models and representations?
by: Jørgensen, Frederik Hytting, et al.
Published: (2025) -
Relational Linear Properties in Language Models: An Empirical Investigation
by: Valer, Giovanni, et al.
Published: (2026) -
Logit Distance Bounds Representational Similarity
by: Nielsen, Beatrix M. G., et al.
Published: (2026)