Towards understanding how attention mechanism works in deep learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ruan, Tianyu, Zhang, Shihua |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Understanding Generalization, Robustness, and Interpretability in Low-Capacity Neural Networks
par: Kumar, Yash
Publié: (2025)
par: Kumar, Yash
Publié: (2025)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
par: Loza, Andrew J., et autres
Publié: (2025)
par: Loza, Andrew J., et autres
Publié: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
par: Mysore, Naveen
Publié: (2026)
par: Mysore, Naveen
Publié: (2026)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
par: Alnemari, Mohammed, et autres
Publié: (2026)
par: Alnemari, Mohammed, et autres
Publié: (2026)
Cross-Architecture Knowledge Distillation (KD) for Retinal Fundus Image Anomaly Detection on NVIDIA Jetson Nano
par: Yilmaz, Berk, et autres
Publié: (2025)
par: Yilmaz, Berk, et autres
Publié: (2025)
The Normalized Difference Layer: A Differentiable Spectral Index Formulation for Deep Learning
par: Lotfi, Ali, et autres
Publié: (2026)
par: Lotfi, Ali, et autres
Publié: (2026)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
par: Dai, Yanning, et autres
Publié: (2026)
par: Dai, Yanning, et autres
Publié: (2026)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
par: Mukherjee, Debdeep, et autres
Publié: (2025)
par: Mukherjee, Debdeep, et autres
Publié: (2025)
TGraphX: Tensor-Aware Graph Neural Network for Multi-Dimensional Feature Learning
par: Sajjadi, Arash, et autres
Publié: (2025)
par: Sajjadi, Arash, et autres
Publié: (2025)
Optimized Gradient Clipping for Noisy Label Learning
par: Ye, Xichen, et autres
Publié: (2024)
par: Ye, Xichen, et autres
Publié: (2024)
Explicit Dropout: Deterministic Regularization for Transformer Architectures
par: Agrawal, Vidhi, et autres
Publié: (2026)
par: Agrawal, Vidhi, et autres
Publié: (2026)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
par: Merin, Aur Shalev
Publié: (2026)
par: Merin, Aur Shalev
Publié: (2026)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
par: Ho, Siu Hang, et autres
Publié: (2025)
par: Ho, Siu Hang, et autres
Publié: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
par: Acuaviva, Pablo, et autres
Publié: (2025)
par: Acuaviva, Pablo, et autres
Publié: (2025)
PolyGLU: State-Conditional Activation Routing in Transformer Feed-Forward Networks
par: Medeiros, Daniel Nobrega
Publié: (2026)
par: Medeiros, Daniel Nobrega
Publié: (2026)
Revisiting GAN with Bayes-Optimal Discrimination
par: Naeini, Mohammadreza Tavasoli, et autres
Publié: (2025)
par: Naeini, Mohammadreza Tavasoli, et autres
Publié: (2025)
Neural Encoding for Image Recall: Human-Like Memory
par: Foussereau, Virgile, et autres
Publié: (2024)
par: Foussereau, Virgile, et autres
Publié: (2024)
JacNet: Learning Functions with Structured Jacobians
par: Lorraine, Jonathan, et autres
Publié: (2024)
par: Lorraine, Jonathan, et autres
Publié: (2024)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
par: Maity, Subhajit, et autres
Publié: (2025)
par: Maity, Subhajit, et autres
Publié: (2025)
Revisiting Non-separable Binary Classification and its Applications in Anomaly Detection
par: Lau, Matthew, et autres
Publié: (2023)
par: Lau, Matthew, et autres
Publié: (2023)
Scalable, Technology-Agnostic Diagnosis and Predictive Maintenance for Point Machine using Deep Learning
par: Di Santi, Eduardo, et autres
Publié: (2025)
par: Di Santi, Eduardo, et autres
Publié: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
par: Qesaraku, Bjorna, et autres
Publié: (2025)
par: Qesaraku, Bjorna, et autres
Publié: (2025)
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
par: Guo, Dongxin, et autres
Publié: (2026)
par: Guo, Dongxin, et autres
Publié: (2026)
A Hybrid Inductive-Transductive Network for Traffic Flow Imputation on Unsampled Locations
par: Rahimiasl, Mohammadmahdi, et autres
Publié: (2025)
par: Rahimiasl, Mohammadmahdi, et autres
Publié: (2025)
Pulse-Driven Neural Architecture: Learnable Oscillatory Dynamics for Robust Continuous-Time Sequence Processing
par: Sharma, Paras
Publié: (2026)
par: Sharma, Paras
Publié: (2026)
GraphNNK -- Graph Classification and Interpretability
par: Bolevic, Zeljko, et autres
Publié: (2026)
par: Bolevic, Zeljko, et autres
Publié: (2026)
Temporal Functional Circuits: From Spline Plots to Faithful Explanations in KAN Forecasting
par: Mysore, Naveen
Publié: (2026)
par: Mysore, Naveen
Publié: (2026)
An in-depth look at approximation via deep and narrow neural networks
par: Dommel, Joris, et autres
Publié: (2025)
par: Dommel, Joris, et autres
Publié: (2025)
Is Cambodia the World's Largest Cashew Producer?
par: Chaya, Veasna, et autres
Publié: (2024)
par: Chaya, Veasna, et autres
Publié: (2024)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
par: Berthier, Louis, et autres
Publié: (2025)
par: Berthier, Louis, et autres
Publié: (2025)
A generalised novel loss function for computational fluid dynamics
par: Cooper-Baldock, Zachary, et autres
Publié: (2024)
par: Cooper-Baldock, Zachary, et autres
Publié: (2024)
H-Model: Dynamic Neural Architectures for Adaptive Processing
par: Hospodarchuk, Dmytro
Publié: (2025)
par: Hospodarchuk, Dmytro
Publié: (2025)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
par: Johnson, David R., et autres
Publié: (2025)
par: Johnson, David R., et autres
Publié: (2025)
Divergence-Based Similarity Function for Multi-View Contrastive Learning
par: Jeon, Jae Hyoung, et autres
Publié: (2025)
par: Jeon, Jae Hyoung, et autres
Publié: (2025)
CellARC: Measuring Intelligence with Cellular Automata
par: Lžičař, Miroslav
Publié: (2025)
par: Lžičař, Miroslav
Publié: (2025)
Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures
par: Mahi, Ishrak Hamim, et autres
Publié: (2026)
par: Mahi, Ishrak Hamim, et autres
Publié: (2026)
On the Equivalence of Regression and Classification
par: Jayadeva, et autres
Publié: (2025)
par: Jayadeva, et autres
Publié: (2025)
SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
par: Guo, Dongxin, et autres
Publié: (2026)
par: Guo, Dongxin, et autres
Publié: (2026)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
par: K, Prasanth K, et autres
Publié: (2025)
par: K, Prasanth K, et autres
Publié: (2025)
From Articles to Canopies: Knowledge-Driven Pseudo-Labelling for Tree Species Classification using LLM Experts
par: Romaszewski, Michał, et autres
Publié: (2026)
par: Romaszewski, Michał, et autres
Publié: (2026)
Documents similaires
-
Understanding Generalization, Robustness, and Interpretability in Low-Capacity Neural Networks
par: Kumar, Yash
Publié: (2025) -
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
par: Loza, Andrew J., et autres
Publié: (2025) -
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
par: Mysore, Naveen
Publié: (2026) -
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
par: Alnemari, Mohammed, et autres
Publié: (2026) -
Cross-Architecture Knowledge Distillation (KD) for Retinal Fundus Image Anomaly Detection on NVIDIA Jetson Nano
par: Yilmaz, Berk, et autres
Publié: (2025)