Analogies between Transformer Layers and Power Method
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chenglong, Altafini, Claudio |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multistability of Self-Attention Dynamics in Transformers
by: Altafini, Claudio
Published: (2025)
by: Altafini, Claudio
Published: (2025)
Gradient Flow Equations for Deep Linear Neural Networks: A Survey from a Network Perspective
by: Wendin, Joel, et al.
Published: (2025)
by: Wendin, Joel, et al.
Published: (2025)
Computing frustration and near-monotonicity in deep neural networks
by: Wendin, Joel, et al.
Published: (2025)
by: Wendin, Joel, et al.
Published: (2025)
Looped Transformers with Layer Normalization Provably Learn the Power Method
by: Wu, Lyumin, et al.
Published: (2026)
by: Wu, Lyumin, et al.
Published: (2026)
Counting in Small Transformers: The Delicate Interplay between Attention and Feed-Forward Layers
by: Behrens, Freya, et al.
Published: (2024)
by: Behrens, Freya, et al.
Published: (2024)
Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning
by: Hellwig, Philipp, et al.
Published: (2026)
by: Hellwig, Philipp, et al.
Published: (2026)
InfoFlow: A Framework for Multi-Layer Transformer Analysis
by: Yu, Penghao, et al.
Published: (2026)
by: Yu, Penghao, et al.
Published: (2026)
AoA-Based Physical Layer Authentication in Analog Arrays under Impersonation Attacks
by: Srinivasan, Muralikrishnan, et al.
Published: (2024)
by: Srinivasan, Muralikrishnan, et al.
Published: (2024)
Matterhorn: Efficient Analog Sparse Spiking Transformer Architecture with Masked Time-To-First-Spike Encoding
by: Yan, Zhanglu, et al.
Published: (2026)
by: Yan, Zhanglu, et al.
Published: (2026)
Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
by: Xu, Ruichen, et al.
Published: (2026)
by: Xu, Ruichen, et al.
Published: (2026)
Predicting Stress-strain Behaviors of Additively Manufactured Materials via Loss-based and Activation-based Physics-informed Machine Learning
by: Duan, Chenglong, et al.
Published: (2026)
by: Duan, Chenglong, et al.
Published: (2026)
A Dynamic Time Warping-Transfer Learning Approach to Transferring Knowledge in Stress-strain Behaviors from Polymers to Metals: An Affordable and Generalizable Additive Manufacturing Part Qualification Framework
by: Duan, Chenglong, et al.
Published: (2025)
by: Duan, Chenglong, et al.
Published: (2025)
Deep Clustering Evaluation: How to Validate Internal Clustering Validation Measures
by: Wang, Zeya, et al.
Published: (2024)
by: Wang, Zeya, et al.
Published: (2024)
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
AnalogFed: Federated Discovery of Analog Circuit Topologies with Generative AI
by: Li, Qiufeng, et al.
Published: (2025)
by: Li, Qiufeng, et al.
Published: (2025)
Analog Foundation Models
by: Büchel, Julian, et al.
Published: (2025)
by: Büchel, Julian, et al.
Published: (2025)
A Regularized Newton Method for Nonconvex Optimization with Global and Local Complexity Guarantees
by: Zhou, Yuhao, et al.
Published: (2025)
by: Zhou, Yuhao, et al.
Published: (2025)
Large Language Models as Analogical Reasoners
by: Yasunaga, Michihiro, et al.
Published: (2023)
by: Yasunaga, Michihiro, et al.
Published: (2023)
Analogy between Boltzmann machines and Feynman path integrals
by: Iyengar, Srinivasan S., et al.
Published: (2023)
by: Iyengar, Srinivasan S., et al.
Published: (2023)
An Analog and Digital Hybrid Attention Accelerator for Transformers with Charge-based In-memory Computing
by: Moradifirouzabadi, Ashkan, et al.
Published: (2024)
by: Moradifirouzabadi, Ashkan, et al.
Published: (2024)
ZeroSim: Zero-Shot Analog Circuit Evaluation with Unified Transformer Embeddings
by: Yang, Xiaomeng, et al.
Published: (2025)
by: Yang, Xiaomeng, et al.
Published: (2025)
ARTEMIS: A Mixed Analog-Stochastic In-DRAM Accelerator for Transformer Neural Networks
by: Afifi, Salma, et al.
Published: (2024)
by: Afifi, Salma, et al.
Published: (2024)
On the Expressive Power of Floating-Point Transformers
by: Park, Sejun, et al.
Published: (2026)
by: Park, Sejun, et al.
Published: (2026)
On the Expressive Power of Contextual Relations in Transformers
by: Fraiman, Demián
Published: (2026)
by: Fraiman, Demián
Published: (2026)
Conformal Transformations for Symmetric Power Transformers
by: Kumar, Saurabh, et al.
Published: (2025)
by: Kumar, Saurabh, et al.
Published: (2025)
On the Expressive Power and Limitations of Multi-Layer SSMs
by: Zubić, Nikola, et al.
Published: (2026)
by: Zubić, Nikola, et al.
Published: (2026)
Decentralized Differentially Private Power Method
by: Campbell, Andrew, et al.
Published: (2025)
by: Campbell, Andrew, et al.
Published: (2025)
A Power Transform
by: Barron, Jonathan T.
Published: (2025)
by: Barron, Jonathan T.
Published: (2025)
AnalogCoder-Pro: Unifying Analog Circuit Generation and Optimization via Multi-modal LLMs
by: Lai, Yao, et al.
Published: (2025)
by: Lai, Yao, et al.
Published: (2025)
SPIRIT: Low Power Seizure Prediction using Unsupervised Online-Learning and Zoom Analog Frontends
by: Pandey, Aviral, et al.
Published: (2024)
by: Pandey, Aviral, et al.
Published: (2024)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
by: Fang, Jiaxun, et al.
Published: (2025)
by: Fang, Jiaxun, et al.
Published: (2025)
Exact Tensor Completion Powered by Slim Transforms
by: Ge, Li, et al.
Published: (2024)
by: Ge, Li, et al.
Published: (2024)
Learning to Skip the Middle Layers of Transformers
by: Lawson, Tim, et al.
Published: (2025)
by: Lawson, Tim, et al.
Published: (2025)
AutoCompress: Critical Layer Isolation for Efficient Transformer Compression
by: Thorat, Archit
Published: (2026)
by: Thorat, Archit
Published: (2026)
GRIT: Graph Transformer For Internal Ice Layer Thickness Prediction
by: Liu, Zesheng, et al.
Published: (2025)
by: Liu, Zesheng, et al.
Published: (2025)
Learning to Select MCP Algorithms: From Traditional ML to Dual-Channel GAT-MLP
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
INSIGHT: Universal Neural Simulator for Analog Circuits Harnessing Autoregressive Transformers
by: Poddar, Souradip, et al.
Published: (2024)
by: Poddar, Souradip, et al.
Published: (2024)
RACE-IT: A Reconfigurable Analog Computing Engine for In-Memory Transformer Acceleration
by: Zhao, Lei, et al.
Published: (2023)
by: Zhao, Lei, et al.
Published: (2023)
A High-accuracy Calibration Method of Transient TSEPs for Power Semiconductor Devices
by: Zhang, Qinghao, et al.
Published: (2025)
by: Zhang, Qinghao, et al.
Published: (2025)
The Power of Second Order Methods for Sequence Preconditioning
by: Marsden, Annie, et al.
Published: (2026)
by: Marsden, Annie, et al.
Published: (2026)
Similar Items
-
Multistability of Self-Attention Dynamics in Transformers
by: Altafini, Claudio
Published: (2025) -
Gradient Flow Equations for Deep Linear Neural Networks: A Survey from a Network Perspective
by: Wendin, Joel, et al.
Published: (2025) -
Computing frustration and near-monotonicity in deep neural networks
by: Wendin, Joel, et al.
Published: (2025) -
Looped Transformers with Layer Normalization Provably Learn the Power Method
by: Wu, Lyumin, et al.
Published: (2026) -
Counting in Small Transformers: The Delicate Interplay between Attention and Feed-Forward Layers
by: Behrens, Freya, et al.
Published: (2024)