Focus and Dilution: The Multi-stage Learning Process of Attention
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Zheng-An, Lin, Pengxiao, Xu, Zhi-Qin John, Luo, Tao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Identity Bridge: Enabling Implicit Reasoning via Shared Latent Memory
por: Lin, Pengxiao, et al.
Publicado: (2025)
por: Lin, Pengxiao, et al.
Publicado: (2025)
Reasoning Bias of Next Token Prediction Training
por: Lin, Pengxiao, et al.
Publicado: (2025)
por: Lin, Pengxiao, et al.
Publicado: (2025)
Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
por: Chen, Tianyi, et al.
Publicado: (2025)
por: Chen, Tianyi, et al.
Publicado: (2025)
Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
por: Zhang, Zhongwang, et al.
Publicado: (2024)
por: Zhang, Zhongwang, et al.
Publicado: (2024)
Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
por: Zhang, Zhongwang, et al.
Publicado: (2025)
por: Zhang, Zhongwang, et al.
Publicado: (2025)
Overview frequency principle/spectral bias in deep learning
por: Xu, Zhi-Qin John, et al.
Publicado: (2022)
por: Xu, Zhi-Qin John, et al.
Publicado: (2022)
Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
por: Bai, Zhiwei, et al.
Publicado: (2022)
por: Bai, Zhiwei, et al.
Publicado: (2022)
Efficient and Flexible Method for Reducing Moderate-size Deep Neural Networks with Condensation
por: Chen, Tianyi, et al.
Publicado: (2024)
por: Chen, Tianyi, et al.
Publicado: (2024)
Filter, Obstruct and Dilute: Defending Against Backdoor Attacks on Semi-Supervised Learning
por: Wang, Xinrui, et al.
Publicado: (2025)
por: Wang, Xinrui, et al.
Publicado: (2025)
Neighbourhood Transformer: Switchable Attention for Monophily-Aware Graph Learning
por: Luo, Yi, et al.
Publicado: (2026)
por: Luo, Yi, et al.
Publicado: (2026)
Focus Where It Matters: Graph Selective State Focused Attention Networks
por: Vashistha, Shikhar, et al.
Publicado: (2024)
por: Vashistha, Shikhar, et al.
Publicado: (2024)
On Multi-Stage Loss Dynamics in Neural Networks: Mechanisms of Plateau and Descent Stages
por: Chen, Zheng-An, et al.
Publicado: (2024)
por: Chen, Zheng-An, et al.
Publicado: (2024)
Learning to Reflect: Hierarchical Multi-Agent Reinforcement Learning for CSI-Free mmWave Beam-Focusing
por: Le, Hieu, et al.
Publicado: (2026)
por: Le, Hieu, et al.
Publicado: (2026)
Diluting Restricted Boltzmann Machines
por: Díaz-Faloh, C., et al.
Publicado: (2025)
por: Díaz-Faloh, C., et al.
Publicado: (2025)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
por: Ram, Dhananjay, et al.
Publicado: (2025)
por: Ram, Dhananjay, et al.
Publicado: (2025)
Urban-Focused Multi-Task Offline Reinforcement Learning with Contrastive Data Sharing
por: Zhao, Xinbo, et al.
Publicado: (2024)
por: Zhao, Xinbo, et al.
Publicado: (2024)
Probability Signature: Bridging Data Semantics and Embedding Structure in Language Models
por: Yao, Junjie, et al.
Publicado: (2025)
por: Yao, Junjie, et al.
Publicado: (2025)
From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
por: Chen, Zheng-An, et al.
Publicado: (2025)
por: Chen, Zheng-An, et al.
Publicado: (2025)
M$^{2}$M: Learning controllable Multi of experts and multi-scale operators are the Partial Differential Equations need
por: Liang, Aoming, et al.
Publicado: (2024)
por: Liang, Aoming, et al.
Publicado: (2024)
Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
por: Xu, Zhi-Qin John, et al.
Publicado: (2019)
por: Xu, Zhi-Qin John, et al.
Publicado: (2019)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
por: Fu, Zichuan, et al.
Publicado: (2026)
por: Fu, Zichuan, et al.
Publicado: (2026)
Decision Focused Causal Learning for Direct Counterfactual Marketing Optimization
por: Zhou, Hao, et al.
Publicado: (2024)
por: Zhou, Hao, et al.
Publicado: (2024)
Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving
por: Zheng, Zhi, et al.
Publicado: (2024)
por: Zheng, Zhi, et al.
Publicado: (2024)
Scalable Complexity Control Facilitates Reasoning Ability of LLMs
por: Hang, Liangkai, et al.
Publicado: (2025)
por: Hang, Liangkai, et al.
Publicado: (2025)
Multi-view Fuzzy Graph Attention Networks for Enhanced Graph Learning
por: Xing, Jinming, et al.
Publicado: (2024)
por: Xing, Jinming, et al.
Publicado: (2024)
Physical Consistency Bridges Heterogeneous Data in Molecular Multi-Task Learning
por: Ren, Yuxuan, et al.
Publicado: (2024)
por: Ren, Yuxuan, et al.
Publicado: (2024)
Bi-Level Decision-Focused Causal Learning for Large-Scale Marketing Optimization: Bridging Observational and Experimental Data
por: Zhang, Shuli, et al.
Publicado: (2025)
por: Zhang, Shuli, et al.
Publicado: (2025)
An overview of condensation phenomenon in deep learning
por: Xu, Zhi-Qin John, et al.
Publicado: (2025)
por: Xu, Zhi-Qin John, et al.
Publicado: (2025)
Loss Spike in Training Neural Networks
por: Li, Xiaolong, et al.
Publicado: (2023)
por: Li, Xiaolong, et al.
Publicado: (2023)
Quantum Graph Attention Network: A Novel Quantum Multi-Head Attention Mechanism for Graph Learning
por: Ning, An, et al.
Publicado: (2025)
por: Ning, An, et al.
Publicado: (2025)
A Deep Learning Framework for Sequence Mining with Bidirectional LSTM and Multi-Scale Attention
por: Yang, Tao, et al.
Publicado: (2025)
por: Yang, Tao, et al.
Publicado: (2025)
Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning
por: Allegue, Daniel De Dios, et al.
Publicado: (2025)
por: Allegue, Daniel De Dios, et al.
Publicado: (2025)
Multi-modal Data based Semi-Supervised Learning for Vehicle Positioning
por: Huan, Ouwen, et al.
Publicado: (2024)
por: Huan, Ouwen, et al.
Publicado: (2024)
MCGM: Multi-stage Clustered Global Modeling for Long-range Interactions in Molecules
por: Pan, Haodong, et al.
Publicado: (2025)
por: Pan, Haodong, et al.
Publicado: (2025)
Focusing Influence Mechanism for Multi-Agent Reinforcement Learning
por: Park, Yisak, et al.
Publicado: (2025)
por: Park, Yisak, et al.
Publicado: (2025)
Multi-Objective Multi-Agent Bandits: From Learning Efficiency to Fairness Optimization
por: Wang, John, et al.
Publicado: (2026)
por: Wang, John, et al.
Publicado: (2026)
Revealing the Learning Process in Reinforcement Learning Agents Through Attention-Oriented Metrics
por: Beylier, Charlotte, et al.
Publicado: (2024)
por: Beylier, Charlotte, et al.
Publicado: (2024)
Learning a Mini-batch Graph Transformer via Two-stage Interaction Augmentation
por: Li, Wenda, et al.
Publicado: (2024)
por: Li, Wenda, et al.
Publicado: (2024)
Temporal-Aware Graph Attention Network for Cryptocurrency Transaction Fraud Detection
por: Zheng, Zhi, et al.
Publicado: (2025)
por: Zheng, Zhi, et al.
Publicado: (2025)
Two-stage Risk Control with Application to Ranked Retrieval
por: Xu, Yunpeng, et al.
Publicado: (2024)
por: Xu, Yunpeng, et al.
Publicado: (2024)
Ejemplares similares
-
Identity Bridge: Enabling Implicit Reasoning via Shared Latent Memory
por: Lin, Pengxiao, et al.
Publicado: (2025) -
Reasoning Bias of Next Token Prediction Training
por: Lin, Pengxiao, et al.
Publicado: (2025) -
Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
por: Chen, Tianyi, et al.
Publicado: (2025) -
Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
por: Zhang, Zhongwang, et al.
Publicado: (2024) -
Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
por: Zhang, Zhongwang, et al.
Publicado: (2025)