InfoFlow: A Framework for Multi-Layer Transformer Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Penghao, Jiang, Haotian, Bao, Zeyu, Li, Qianxiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Effect of Attention Head Count on Transformer Approximation
by: Yu, Penghao, et al.
Published: (2025)
by: Yu, Penghao, et al.
Published: (2025)
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
by: Bao, Zeyu, et al.
Published: (2025)
by: Bao, Zeyu, et al.
Published: (2025)
InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context
by: Teng, Xin, et al.
Published: (2026)
by: Teng, Xin, et al.
Published: (2026)
Allocation of Parameters in Transformers
by: Yu, Ruoxi, et al.
Published: (2025)
by: Yu, Ruoxi, et al.
Published: (2025)
Approximation Rate of the Transformer Architecture for Sequence Modeling
by: Jiang, Haotian, et al.
Published: (2023)
by: Jiang, Haotian, et al.
Published: (2023)
Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions
by: Jiang, Haotian, et al.
Published: (2025)
by: Jiang, Haotian, et al.
Published: (2025)
InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
by: Luo, Kun, et al.
Published: (2025)
by: Luo, Kun, et al.
Published: (2025)
From Generalization Analysis to Optimization Designs for State Space Models
by: Liu, Fusheng, et al.
Published: (2024)
by: Liu, Fusheng, et al.
Published: (2024)
Learning task-specific predictive models for scientific computing
by: Yin, Jianyuan, et al.
Published: (2025)
by: Yin, Jianyuan, et al.
Published: (2025)
Autocorrelation Matters: Understanding the Role of Initialization Schemes for State Space Models
by: Liu, Fusheng, et al.
Published: (2024)
by: Liu, Fusheng, et al.
Published: (2024)
Accelerating Legacy Numerical Solvers by Non-intrusive Gradient-based Meta-solving
by: Arisaka, Sohei, et al.
Published: (2024)
by: Arisaka, Sohei, et al.
Published: (2024)
Unifying back-propagation and forward-forward algorithms through model predictive control
by: Ren, Lianhai, et al.
Published: (2024)
by: Ren, Lianhai, et al.
Published: (2024)
DynGMA: a robust approach for learning stochastic differential equations from data
by: Zhu, Aiqing, et al.
Published: (2024)
by: Zhu, Aiqing, et al.
Published: (2024)
Learning Macroscopic Dynamics from Partial Microscopic Observations
by: Chen, Mengyi, et al.
Published: (2024)
by: Chen, Mengyi, et al.
Published: (2024)
InfoQ: Mixed-Precision Quantization via Global Information Flow
by: Akbulut, Mehmet Emre, et al.
Published: (2025)
by: Akbulut, Mehmet Emre, et al.
Published: (2025)
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
by: Wang, Shida, et al.
Published: (2023)
by: Wang, Shida, et al.
Published: (2023)
Learning Permutation-invariant Macroscopic Dynamics
by: Han, Zhichao, et al.
Published: (2026)
by: Han, Zhichao, et al.
Published: (2026)
InfoDPCCA: Information-Theoretic Dynamic Probabilistic Canonical Correlation Analysis
by: Tang, Shiqin, et al.
Published: (2025)
by: Tang, Shiqin, et al.
Published: (2025)
SG-OIF: A Stability-Guided Online Influence Framework for Reliable Vision Data
by: Rao, Penghao, et al.
Published: (2025)
by: Rao, Penghao, et al.
Published: (2025)
Machine Unlearning under Retain-Forget Entanglement
by: Cheng, Jingpu, et al.
Published: (2026)
by: Cheng, Jingpu, et al.
Published: (2026)
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
by: Wang, Shida, et al.
Published: (2023)
by: Wang, Shida, et al.
Published: (2023)
InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
by: Wang, Youjin, et al.
Published: (2026)
by: Wang, Youjin, et al.
Published: (2026)
Info-Coevolution: An Efficient Framework for Data Model Coevolution
by: Qin, Ziheng, et al.
Published: (2025)
by: Qin, Ziheng, et al.
Published: (2025)
A unified framework for establishing the universal approximation of transformer-type architectures
by: Cheng, Jingpu, et al.
Published: (2025)
by: Cheng, Jingpu, et al.
Published: (2025)
Continuity-Preserving Convolutional Autoencoders for Learning Continuous Latent Dynamical Models from Images
by: Zhu, Aiqing, et al.
Published: (2025)
by: Zhu, Aiqing, et al.
Published: (2025)
NetInfoF Framework: Measuring and Exploiting Network Usable Information
by: Lee, Meng-Chieh, et al.
Published: (2024)
by: Lee, Meng-Chieh, et al.
Published: (2024)
Deep learning and the rate of approximation by flows
by: Cheng, Jingpu, et al.
Published: (2026)
by: Cheng, Jingpu, et al.
Published: (2026)
Scalable learning of macroscopic stochastic dynamics
by: Chen, Mengyi, et al.
Published: (2025)
by: Chen, Mengyi, et al.
Published: (2025)
Identifiable learning of dissipative dynamics
by: Zhu, Aiqing, et al.
Published: (2025)
by: Zhu, Aiqing, et al.
Published: (2025)
Double InfoGAN for Contrastive Analysis
by: Carton, Florence, et al.
Published: (2024)
by: Carton, Florence, et al.
Published: (2024)
Mitigating distribution shift in machine learning-augmented hybrid simulation
by: Zhao, Jiaxi, et al.
Published: (2024)
by: Zhao, Jiaxi, et al.
Published: (2024)
$f$-MICL: Understanding and Generalizing InfoNCE-based Contrastive Learning
by: Lu, Yiwei, et al.
Published: (2024)
by: Lu, Yiwei, et al.
Published: (2024)
Info-CELS: Informative Saliency Map Guided Counterfactual Explanation
by: Li, Peiyu, et al.
Published: (2024)
by: Li, Peiyu, et al.
Published: (2024)
Flowing Through Layers: A Continuous Dynamical Systems Perspective on Transformers
by: Fein-Ashley, Jacob
Published: (2025)
by: Fein-Ashley, Jacob
Published: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
by: Miao, Yuchun, et al.
Published: (2024)
by: Miao, Yuchun, et al.
Published: (2024)
InfoPos: A Design Support Framework for ML-Assisted Fault Detection and Identification in Industrial Cyber-Physical Systems
by: Odyurt, Uraz, et al.
Published: (2025)
by: Odyurt, Uraz, et al.
Published: (2025)
Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
by: Wang, Penghao, et al.
Published: (2025)
by: Wang, Penghao, et al.
Published: (2025)
On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
by: Li, Zhong, et al.
Published: (2020)
by: Li, Zhong, et al.
Published: (2020)
Terminally constrained flow-based generative models from an optimal control perspective
by: Gao, Weiguo, et al.
Published: (2026)
by: Gao, Weiguo, et al.
Published: (2026)
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
by: Song, Zhuo-Yang, et al.
Published: (2025)
by: Song, Zhuo-Yang, et al.
Published: (2025)
Similar Items
-
The Effect of Attention Head Count on Transformer Approximation
by: Yu, Penghao, et al.
Published: (2025) -
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
by: Bao, Zeyu, et al.
Published: (2025) -
InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context
by: Teng, Xin, et al.
Published: (2026) -
Allocation of Parameters in Transformers
by: Yu, Ruoxi, et al.
Published: (2025) -
Approximation Rate of the Transformer Architecture for Sequence Modeling
by: Jiang, Haotian, et al.
Published: (2023)