Geometric Signatures of Compositionality Across a Language Model's Lifetime
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jin Hwa, Jiralerspong, Thomas, Yu, Lei, Bengio, Yoshua, Cheng, Emily |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
Efficient Causal Graph Discovery Using Large Language Models
by: Jiralerspong, Thomas, et al.
Published: (2024)
by: Jiralerspong, Thomas, et al.
Published: (2024)
Shaping Inductive Bias in Diffusion Models through Frequency-Based Noise Control
by: Jiralerspong, Thomas, et al.
Published: (2025)
by: Jiralerspong, Thomas, et al.
Published: (2025)
Language Models Can Reduce Asymmetry in Information Markets
by: Rahaman, Nasim, et al.
Published: (2024)
by: Rahaman, Nasim, et al.
Published: (2024)
A Survey on Large Language Models from Concept to Implementation
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
Learning What Matters: Steering Diffusion via Spectrally Anisotropic Forward Noise
by: Scimeca, Luca, et al.
Published: (2025)
by: Scimeca, Luca, et al.
Published: (2025)
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023)
by: Delétang, Grégoire, et al.
Published: (2023)
The Information of Large Language Model Geometry
by: Tan, Zhiquan, et al.
Published: (2024)
by: Tan, Zhiquan, et al.
Published: (2024)
Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models
by: Wei, Lai, et al.
Published: (2024)
by: Wei, Lai, et al.
Published: (2024)
Memorization-Compression Cycles Improve Generalization
by: Yu, Fangyuan
Published: (2025)
by: Yu, Fangyuan
Published: (2025)
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
by: Garg, Nikhil, et al.
Published: (2026)
by: Garg, Nikhil, et al.
Published: (2026)
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
by: Badger, Benjamin L., et al.
Published: (2025)
by: Badger, Benjamin L., et al.
Published: (2025)
The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
by: Wang, Hanyang, et al.
Published: (2026)
by: Wang, Hanyang, et al.
Published: (2026)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025)
by: Kang, Justin Singh, et al.
Published: (2025)
L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling
by: Chen, Zhuo, et al.
Published: (2025)
by: Chen, Zhuo, et al.
Published: (2025)
A Training-free Method for LLM Text Attribution
by: Radvand, Tara, et al.
Published: (2025)
by: Radvand, Tara, et al.
Published: (2025)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
by: Anwar, Usman, et al.
Published: (2026)
by: Anwar, Usman, et al.
Published: (2026)
Optimal Quantization for Matrix Multiplication
by: Ordentlich, Or, et al.
Published: (2024)
by: Ordentlich, Or, et al.
Published: (2024)
Compression Represents Intelligence Linearly
by: Huang, Yuzhen, et al.
Published: (2024)
by: Huang, Yuzhen, et al.
Published: (2024)
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
by: Krasnovsky, Anatoly A.
Published: (2025)
by: Krasnovsky, Anatoly A.
Published: (2025)
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability
by: Omidvar, Hamed, et al.
Published: (2026)
by: Omidvar, Hamed, et al.
Published: (2026)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
by: Wieser, Frederico, et al.
Published: (2025)
by: Wieser, Frederico, et al.
Published: (2025)
SQuat: Subspace-orthogonal KV Cache Quantization
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
An Information Theoretic Perspective on Agentic System Design
by: He, Shizhe, et al.
Published: (2025)
by: He, Shizhe, et al.
Published: (2025)
The Causal Description Gap: Information-Theoretic Separations Across Pearl's Hierarchy
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
by: Lee, Jin Hwa, et al.
Published: (2025)
by: Lee, Jin Hwa, et al.
Published: (2025)
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
by: Lo, Chung-Hsiang, et al.
Published: (2026)
by: Lo, Chung-Hsiang, et al.
Published: (2026)
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
by: Xu, Shizhou, et al.
Published: (2025)
by: Xu, Shizhou, et al.
Published: (2025)
TableRAG: Million-Token Table Understanding with Language Models
by: Chen, Si-An, et al.
Published: (2024)
by: Chen, Si-An, et al.
Published: (2024)
Estimating Mutual Information between Time Series and Temporal Event Sequences Across Diverse Analysis Tasks
by: Hu, Haoji, et al.
Published: (2026)
by: Hu, Haoji, et al.
Published: (2026)
Large Language Models for Telecom: Forthcoming Impact on the Industry
by: Maatouk, Ali, et al.
Published: (2023)
by: Maatouk, Ali, et al.
Published: (2023)
The Factuality of Large Language Models in the Legal Domain
by: Hamdani, Rajaa El, et al.
Published: (2024)
by: Hamdani, Rajaa El, et al.
Published: (2024)
Confidence-Based Decoding is Provably Efficient for Diffusion Language Models
by: Cai, Changxiao, et al.
Published: (2026)
by: Cai, Changxiao, et al.
Published: (2026)
On the Limits of Self-Improving in Large Language Models: The Singularity Is Not Near Without Symbolic Model Synthesis
by: Zenil, Hector
Published: (2026)
by: Zenil, Hector
Published: (2026)
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
by: Razzhigaev, Anton, et al.
Published: (2023)
by: Razzhigaev, Anton, et al.
Published: (2023)
WDMoE: Wireless Distributed Large Language Models with Mixture of Experts
by: Xue, Nan, et al.
Published: (2024)
by: Xue, Nan, et al.
Published: (2024)
Similar Items
-
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024) -
Efficient Causal Graph Discovery Using Large Language Models
by: Jiralerspong, Thomas, et al.
Published: (2024) -
Shaping Inductive Bias in Diffusion Models through Frequency-Based Noise Control
by: Jiralerspong, Thomas, et al.
Published: (2025) -
Language Models Can Reduce Asymmetry in Information Markets
by: Rahaman, Nasim, et al.
Published: (2024) -
A Survey on Large Language Models from Concept to Implementation
by: Wang, Chen, et al.
Published: (2024)