Modeling Transformers as complex networks to analyze learning dynamics
Fuente:
arXiv
Saved in:
| Main Author: | Rocchetti, Elisabetta |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unveiling Transformer Perception by Exploring Input Manifolds
by: Benfenati, Alessandro, et al.
Published: (2024)
by: Benfenati, Alessandro, et al.
Published: (2024)
Parallel BiLSTM-Transformer networks for forecasting chaotic dynamics
by: Ma, Junwen, et al.
Published: (2025)
by: Ma, Junwen, et al.
Published: (2025)
How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism
by: Rocchetti, Elisabetta, et al.
Published: (2026)
by: Rocchetti, Elisabetta, et al.
Published: (2026)
Enhanced Transformer architecture for in-context learning of dynamical systems
by: Rufolo, Matteo, et al.
Published: (2024)
by: Rufolo, Matteo, et al.
Published: (2024)
Decoding complexity: how machine learning is redefining scientific discovery
by: Vinuesa, Ricardo, et al.
Published: (2024)
by: Vinuesa, Ricardo, et al.
Published: (2024)
Understanding the dynamics of the frequency bias in neural networks
by: Molina, Juan, et al.
Published: (2024)
by: Molina, Juan, et al.
Published: (2024)
In value-based deep reinforcement learning, a pruned network is a good network
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Gradient-free variational learning with conditional mixture networks
by: Heins, Conor, et al.
Published: (2024)
by: Heins, Conor, et al.
Published: (2024)
How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis
by: Rocchetti, Elisabetta, et al.
Published: (2025)
by: Rocchetti, Elisabetta, et al.
Published: (2025)
Understanding the learned look-ahead behavior of chess neural networks
by: Cruz, Diogo
Published: (2025)
by: Cruz, Diogo
Published: (2025)
Adaptive multiple optimal learning factors for neural network training
by: Challagundla, Jeshwanth
Published: (2024)
by: Challagundla, Jeshwanth
Published: (2024)
Investigating potential causes of Sepsis with Bayesian network structure learning
by: Petrungaro, Bruno, et al.
Published: (2024)
by: Petrungaro, Bruno, et al.
Published: (2024)
A representational framework for learning and encoding structurally enriched trajectories in complex agent environments
by: Catarau-Cotutiu, Corina, et al.
Published: (2025)
by: Catarau-Cotutiu, Corina, et al.
Published: (2025)
Context selectivity with dynamic availability enables lifelong continual learning
by: Barry, Martin, et al.
Published: (2023)
by: Barry, Martin, et al.
Published: (2023)
Feature contamination: Neural networks learn uncorrelated features and fail to generalize
by: Zhang, Tianren, et al.
Published: (2024)
by: Zhang, Tianren, et al.
Published: (2024)
Neural operators struggle to learn complex PDEs in pedestrian mobility: Hughes model case study
by: Chauhan, Prajwal, et al.
Published: (2025)
by: Chauhan, Prajwal, et al.
Published: (2025)
Training instability in deep learning follows low-dimensional dynamical principles
by: Zhang, Zhipeng, et al.
Published: (2026)
by: Zhang, Zhipeng, et al.
Published: (2026)
Decoding the mechanisms of the Hattrick football manager game using Bayesian network structure learning
by: Constantinou, Anthony C., et al.
Published: (2025)
by: Constantinou, Anthony C., et al.
Published: (2025)
FedMSE: Semi-supervised federated learning approach for IoT network intrusion detection
by: Nguyen, Van Tuan, et al.
Published: (2024)
by: Nguyen, Van Tuan, et al.
Published: (2024)
Designing an efficient and equitable humanitarian supply chain dynamically via reinforcement learning
by: Jin, Weijia
Published: (2025)
by: Jin, Weijia
Published: (2025)
Scientific machine learning in ecological systems: A study on the predator-prey dynamics
by: Devgupta, Ranabir, et al.
Published: (2024)
by: Devgupta, Ranabir, et al.
Published: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Representations learnt by SGD and Adaptive learning rules: Conditions that vary sparsity and selectivity in neural networks
by: Park, Jin Hyun
Published: (2022)
by: Park, Jin Hyun
Published: (2022)
LLMs learn governing principles of dynamical systems, revealing an in-context neural scaling law
by: Liu, Toni J. B., et al.
Published: (2024)
by: Liu, Toni J. B., et al.
Published: (2024)
The sample complexity of multi-distribution learning
by: Peng, Binghui
Published: (2023)
by: Peng, Binghui
Published: (2023)
Task complexity shapes internal representations and robustness in neural networks
by: Jankowski, Robert, et al.
Published: (2025)
by: Jankowski, Robert, et al.
Published: (2025)
Zero-shot Meta-learning for Tabular Prediction Tasks with Adversarially Pre-trained Transformer
by: Wu, Yulun, et al.
Published: (2025)
by: Wu, Yulun, et al.
Published: (2025)
Traj-Transformer: Diffusion Models with Transformer for GPS Trajectory Generation
by: Zhang, Zhiyang, et al.
Published: (2025)
by: Zhang, Zhiyang, et al.
Published: (2025)
Target noise: A pre-training based neural network initialization for efficient high resolution learning
by: Wang, Shaowen, et al.
Published: (2026)
by: Wang, Shaowen, et al.
Published: (2026)
An evolutionary perspective on modes of learning in Transformers
by: Ku, Alexander Y., et al.
Published: (2025)
by: Ku, Alexander Y., et al.
Published: (2025)
Boosting long-term forecasting performance for continuous-time dynamic graph networks via data augmentation
by: Tian, Yuxing, et al.
Published: (2023)
by: Tian, Yuxing, et al.
Published: (2023)
Noise Stability of Transformer Models
by: Haris, Themistoklis, et al.
Published: (2026)
by: Haris, Themistoklis, et al.
Published: (2026)
Transformers trained on proteins can learn to attend to Euclidean distance
by: Ellmen, Isaac, et al.
Published: (2025)
by: Ellmen, Isaac, et al.
Published: (2025)
Transforming Multimodal Models into Action Models for Radiotherapy
by: Ferrante, Matteo, et al.
Published: (2025)
by: Ferrante, Matteo, et al.
Published: (2025)
On-site estimation of battery electrochemical parameters via transfer learning based physics-informed neural network approach
by: Yeregui, Josu, et al.
Published: (2025)
by: Yeregui, Josu, et al.
Published: (2025)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
by: Boix-Adsera, Enric, et al.
Published: (2025)
by: Boix-Adsera, Enric, et al.
Published: (2025)
Active management of battery degradation in wireless sensor network using deep reinforcement learning for group battery replacement
by: Jeong, Jong-Hyun, et al.
Published: (2025)
by: Jeong, Jong-Hyun, et al.
Published: (2025)
Deep-layer limit and stability analysis of the basic forward-backward-splitting induced network (II): learning problems
by: Lin, Xuan, et al.
Published: (2026)
by: Lin, Xuan, et al.
Published: (2026)
Data-driven simulator of multi-animal behavior with unknown dynamics via offline and online reinforcement learning
by: Fujii, Keisuke, et al.
Published: (2025)
by: Fujii, Keisuke, et al.
Published: (2025)
The power and limitations of learning quantum dynamics incoherently
by: Jerbi, Sofiene, et al.
Published: (2023)
by: Jerbi, Sofiene, et al.
Published: (2023)
Similar Items
-
Unveiling Transformer Perception by Exploring Input Manifolds
by: Benfenati, Alessandro, et al.
Published: (2024) -
Parallel BiLSTM-Transformer networks for forecasting chaotic dynamics
by: Ma, Junwen, et al.
Published: (2025) -
How LLMs Follow Instructions: Skillful Coordination, Not a Universal Mechanism
by: Rocchetti, Elisabetta, et al.
Published: (2026) -
Enhanced Transformer architecture for in-context learning of dynamical systems
by: Rufolo, Matteo, et al.
Published: (2024) -
Decoding complexity: how machine learning is redefining scientific discovery
by: Vinuesa, Ricardo, et al.
Published: (2024)