Subjective Depth and Timescale Transformers: Learning Where and When to Compute
Fuente:
arXiv
Saved in:
| Main Authors: | Wieser, Frederico, Benfeghoul, Martin, Ammar, Haitham Bou, Wang, Jun, Fountas, Zafeirios |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
by: Oomerjee, Adnan, et al.
Published: (2025)
by: Oomerjee, Adnan, et al.
Published: (2025)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
by: Benfeghoul, Martin, et al.
Published: (2025)
by: Benfeghoul, Martin, et al.
Published: (2025)
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
by: Hazard, Hugo, et al.
Published: (2025)
by: Hazard, Hugo, et al.
Published: (2025)
Human-inspired Episodic Memory for Infinite Context LLMs
by: Fountas, Zafeirios, et al.
Published: (2024)
by: Fountas, Zafeirios, et al.
Published: (2024)
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
by: Fountas, Zafeirios, et al.
Published: (2026)
by: Fountas, Zafeirios, et al.
Published: (2026)
When in Doubt, Think Slow: Iterative Reasoning with Latent Imagination
by: Benfeghoul, Martin, et al.
Published: (2024)
by: Benfeghoul, Martin, et al.
Published: (2024)
Emergent Bayesian Behaviour and Optimal Cue Combination in LLMs
by: Ma, Julian, et al.
Published: (2025)
by: Ma, Julian, et al.
Published: (2025)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
by: Christopoulou, Fenia, et al.
Published: (2024)
by: Christopoulou, Fenia, et al.
Published: (2024)
Mixture of Attentions For Speculative Decoding
by: Zimmer, Matthieu, et al.
Published: (2024)
by: Zimmer, Matthieu, et al.
Published: (2024)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
by: Krasnovsky, Anatoly A.
Published: (2025)
by: Krasnovsky, Anatoly A.
Published: (2025)
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
by: Razzhigaev, Anton, et al.
Published: (2023)
by: Razzhigaev, Anton, et al.
Published: (2023)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
by: Ji, Xiaotong, et al.
Published: (2026)
by: Ji, Xiaotong, et al.
Published: (2026)
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
by: Ji, Xiaotong, et al.
Published: (2026)
by: Ji, Xiaotong, et al.
Published: (2026)
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
A Training-free Method for LLM Text Attribution
by: Radvand, Tara, et al.
Published: (2025)
by: Radvand, Tara, et al.
Published: (2025)
Know Your Limits: Entropy Estimation Modeling for Compression and Generalization
by: Badger, Benjamin L., et al.
Published: (2025)
by: Badger, Benjamin L., et al.
Published: (2025)
Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
by: Anwar, Usman, et al.
Published: (2026)
by: Anwar, Usman, et al.
Published: (2026)
Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models
by: Wei, Lai, et al.
Published: (2024)
by: Wei, Lai, et al.
Published: (2024)
Optimal Quantization for Matrix Multiplication
by: Ordentlich, Or, et al.
Published: (2024)
by: Ordentlich, Or, et al.
Published: (2024)
Language Modeling Is Compression
by: Delétang, Grégoire, et al.
Published: (2023)
by: Delétang, Grégoire, et al.
Published: (2023)
Memorization-Compression Cycles Improve Generalization
by: Yu, Fangyuan
Published: (2025)
by: Yu, Fangyuan
Published: (2025)
Compression Represents Intelligence Linearly
by: Huang, Yuzhen, et al.
Published: (2024)
by: Huang, Yuzhen, et al.
Published: (2024)
A Survey on Large Language Models from Concept to Implementation
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025)
by: Kang, Justin Singh, et al.
Published: (2025)
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability
by: Omidvar, Hamed, et al.
Published: (2026)
by: Omidvar, Hamed, et al.
Published: (2026)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
SQuat: Subspace-orthogonal KV Cache Quantization
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
by: Wang, Hanyang, et al.
Published: (2026)
by: Wang, Hanyang, et al.
Published: (2026)
An Information Theoretic Perspective on Agentic System Design
by: He, Shizhe, et al.
Published: (2025)
by: He, Shizhe, et al.
Published: (2025)
The Information of Large Language Model Geometry
by: Tan, Zhiquan, et al.
Published: (2024)
by: Tan, Zhiquan, et al.
Published: (2024)
Geometric Signatures of Compositionality Across a Language Model's Lifetime
by: Lee, Jin Hwa, et al.
Published: (2024)
by: Lee, Jin Hwa, et al.
Published: (2024)
In-Context Learning for MIMO Equalization Using Transformer-Based Sequence Models
by: Zecchin, Matteo, et al.
Published: (2023)
by: Zecchin, Matteo, et al.
Published: (2023)
Compute-Update Federated Learning: A Lattice Coding Approach Over-the-Air
by: Azimi-Abarghouyi, Seyed Mohammad, et al.
Published: (2024)
by: Azimi-Abarghouyi, Seyed Mohammad, et al.
Published: (2024)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
by: Nguyen, Tu, et al.
Published: (2026)
by: Nguyen, Tu, et al.
Published: (2026)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
by: Roy, Amartya, et al.
Published: (2026)
by: Roy, Amartya, et al.
Published: (2026)
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Similar Items
-
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
by: Oomerjee, Adnan, et al.
Published: (2025) -
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
by: Benfeghoul, Martin, et al.
Published: (2025) -
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
by: Hazard, Hugo, et al.
Published: (2025) -
Human-inspired Episodic Memory for Infinite Context LLMs
by: Fountas, Zafeirios, et al.
Published: (2024) -
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
by: Fountas, Zafeirios, et al.
Published: (2026)