Low-Dimensional Execution Manifolds in Transformer Learning Dynamics: Evidence from Modular Arithmetic Tasks
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Xu, Yongzhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
von: Zhao, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xingyu, et al.
Veröffentlicht: (2026)
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
von: Avinash, Mynampati Sri Ranganadha
Veröffentlicht: (2026)
von: Avinash, Mynampati Sri Ranganadha
Veröffentlicht: (2026)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
von: Hedar, Abdel-Rahman, et al.
Veröffentlicht: (2024)
von: Hedar, Abdel-Rahman, et al.
Veröffentlicht: (2024)
Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks
von: Arora, Rushiv
Veröffentlicht: (2025)
von: Arora, Rushiv
Veröffentlicht: (2025)
DYNAMITE: Dynamic Interplay of Mini-Batch Size and Aggregation Frequency for Federated Learning with Static and Streaming Dataset
von: Liu, Weijie, et al.
Veröffentlicht: (2023)
von: Liu, Weijie, et al.
Veröffentlicht: (2023)
A Simple Generalisation of the Implicit Dynamics of In-Context Learning
von: Innocenti, Francesco, et al.
Veröffentlicht: (2025)
von: Innocenti, Francesco, et al.
Veröffentlicht: (2025)
A Geometric Perspective for High-Dimensional Multiplex Graphs
von: Abdous, Kamel, et al.
Veröffentlicht: (2025)
von: Abdous, Kamel, et al.
Veröffentlicht: (2025)
Not All Transitions Matter: Evidence from PPO
von: Basnet, Ajhesh
Veröffentlicht: (2026)
von: Basnet, Ajhesh
Veröffentlicht: (2026)
SourceSplice: Source Selection for Machine Learning Tasks
von: Singh, Ambarish, et al.
Veröffentlicht: (2025)
von: Singh, Ambarish, et al.
Veröffentlicht: (2025)
Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent Space
von: Chen, Zitao, et al.
Veröffentlicht: (2025)
von: Chen, Zitao, et al.
Veröffentlicht: (2025)
Path-Coupled Bellman Flows for Distributional Reinforcement Learning
von: Xu, Boyang, et al.
Veröffentlicht: (2026)
von: Xu, Boyang, et al.
Veröffentlicht: (2026)
Evidential Deep Active Learning for Semi-Supervised Classification
von: Zhao, Shenkai, et al.
Veröffentlicht: (2025)
von: Zhao, Shenkai, et al.
Veröffentlicht: (2025)
Behavior Learning (BL): Learning Hierarchical Optimization Structures from Data
von: Ma, Zhenyao, et al.
Veröffentlicht: (2026)
von: Ma, Zhenyao, et al.
Veröffentlicht: (2026)
Superposition Is Not Necessary: A Mechanistic Interpretability Analysis of Transformer Representations for Time Series Forecasting
von: Yıldırım, Alper
Veröffentlicht: (2026)
von: Yıldırım, Alper
Veröffentlicht: (2026)
One Router to Route Them All: Homogeneous Expert Routing for Heterogeneous Graph Transformers
von: Shakirov, Georgiy, et al.
Veröffentlicht: (2025)
von: Shakirov, Georgiy, et al.
Veröffentlicht: (2025)
TACO: Tackling Over-correction in Federated Learning with Tailored Adaptive Correction
von: Liu, Weijie, et al.
Veröffentlicht: (2025)
von: Liu, Weijie, et al.
Veröffentlicht: (2025)
DRFormer: Multi-Scale Transformer Utilizing Diverse Receptive Fields for Long Time-Series Forecasting
von: Ding, Ruixin, et al.
Veröffentlicht: (2024)
von: Ding, Ruixin, et al.
Veröffentlicht: (2024)
Dynamics Reveals Structure: Challenging the Linear Propagation Assumption
von: Chang, Hoyeon, et al.
Veröffentlicht: (2026)
von: Chang, Hoyeon, et al.
Veröffentlicht: (2026)
TS-ACL: Closed-Form Solution for Time Series-oriented Continual Learning
von: Li, Jiaxu, et al.
Veröffentlicht: (2024)
von: Li, Jiaxu, et al.
Veröffentlicht: (2024)
Machine Learning vs Deep Learning: The Generalization Problem
von: Bay, Yong Yi, et al.
Veröffentlicht: (2024)
von: Bay, Yong Yi, et al.
Veröffentlicht: (2024)
Expressive Value Learning for Scalable Offline Reinforcement Learning
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
von: Espinosa-Dice, Nicolas, et al.
Veröffentlicht: (2025)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
von: Vincze, Mátyás, et al.
Veröffentlicht: (2024)
von: Vincze, Mátyás, et al.
Veröffentlicht: (2024)
EEG Sleep Stage Classification with Continuous Wavelet Transform and Deep Learning
von: Gashti, Mehdi Zekriyapanah, et al.
Veröffentlicht: (2025)
von: Gashti, Mehdi Zekriyapanah, et al.
Veröffentlicht: (2025)
Bounded Ratio Reinforcement Learning
von: Ao, Yunke, et al.
Veröffentlicht: (2026)
von: Ao, Yunke, et al.
Veröffentlicht: (2026)
Why Online Reinforcement Learning is Causal
von: Schulte, Oliver, et al.
Veröffentlicht: (2024)
von: Schulte, Oliver, et al.
Veröffentlicht: (2024)
CircuitProbe: Predicting Reasoning Circuits in Transformers via Stability Zone Detection
von: Panuganti, Rajkiran
Veröffentlicht: (2026)
von: Panuganti, Rajkiran
Veröffentlicht: (2026)
Understanding Goal Generalisation in Sequential Reinforcement Learning
von: Brown, Jason Ross, et al.
Veröffentlicht: (2026)
von: Brown, Jason Ross, et al.
Veröffentlicht: (2026)
DataRater: Meta-Learned Dataset Curation
von: Calian, Dan A., et al.
Veröffentlicht: (2025)
von: Calian, Dan A., et al.
Veröffentlicht: (2025)
Extrinsicaly Rewarded Soft Q Imitation Learning with Discriminator
von: Furuyama, Ryoma, et al.
Veröffentlicht: (2024)
von: Furuyama, Ryoma, et al.
Veröffentlicht: (2024)
Less is More: Learning Graph Tasks with Just LLMs
von: Shirai, Sola, et al.
Veröffentlicht: (2025)
von: Shirai, Sola, et al.
Veröffentlicht: (2025)
Integrating Causality with Neurochaos Learning: Proposed Approach and Research Agenda
von: Narendra, Nanjangud C., et al.
Veröffentlicht: (2025)
von: Narendra, Nanjangud C., et al.
Veröffentlicht: (2025)
The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models
von: Liu, Ming
Veröffentlicht: (2026)
von: Liu, Ming
Veröffentlicht: (2026)
Load and Renewable Energy Forecasting Using Deep Learning for Grid Stability
von: Sarkar, Kamal
Veröffentlicht: (2025)
von: Sarkar, Kamal
Veröffentlicht: (2025)
Learning Agents With Prioritization and Parameter Noise in Continuous State and Action Space
von: Mangannavar, Rajesh, et al.
Veröffentlicht: (2024)
von: Mangannavar, Rajesh, et al.
Veröffentlicht: (2024)
What changes after deployment? A survey on On-device Learning in TinyML
von: Pavan, Massimo, et al.
Veröffentlicht: (2026)
von: Pavan, Massimo, et al.
Veröffentlicht: (2026)
FreRA: A Frequency-Refined Augmentation for Contrastive Learning on Time Series Classification
von: Tian, Tian, et al.
Veröffentlicht: (2025)
von: Tian, Tian, et al.
Veröffentlicht: (2025)
A Practical Approach to using Supervised Machine Learning Models to Classify Aviation Safety Occurrences
von: Siow, Bryan Y.
Veröffentlicht: (2025)
von: Siow, Bryan Y.
Veröffentlicht: (2025)
DELTA: Variational Disentangled Learning for Privacy-Preserving Data Reprogramming
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2025)
von: Malarkkan, Arun Vignesh, et al.
Veröffentlicht: (2025)
What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators
von: Zhang, Xinyu
Veröffentlicht: (2026)
von: Zhang, Xinyu
Veröffentlicht: (2026)
Working Paper: Active Causal Structure Learning with Latent Variables: Towards Learning to Detour in Autonomous Robots
von: Riscos, Pablo de los, et al.
Veröffentlicht: (2024)
von: Riscos, Pablo de los, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
von: Zhao, Xingyu, et al.
Veröffentlicht: (2026) -
Task-Conditioned Routing Signatures in Sparse Mixture-of-Experts Transformers
von: Avinash, Mynampati Sri Ranganadha
Veröffentlicht: (2026) -
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
von: Hedar, Abdel-Rahman, et al.
Veröffentlicht: (2024) -
Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks
von: Arora, Rushiv
Veröffentlicht: (2025) -
DYNAMITE: Dynamic Interplay of Mini-Batch Size and Aggregation Frequency for Federated Learning with Static and Streaming Dataset
von: Liu, Weijie, et al.
Veröffentlicht: (2023)