Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pham, Duy-Tung, Nguyen, An The, Tran, Viet-Hoang, Chung, Nhan-Phu, Tong, Xin T., Nguyen, Tan M., Vo, Thieu N. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Demystifying the Token Dynamics of Deep Selective State Space Models
von: Vo, Thieu N, et al.
Veröffentlicht: (2024)
von: Vo, Thieu N, et al.
Veröffentlicht: (2024)
Equivariant Polynomial Functional Networks
von: Vo, Thieu N., et al.
Veröffentlicht: (2024)
von: Vo, Thieu N., et al.
Veröffentlicht: (2024)
Equivariant Neural Functional Networks for Transformers
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024)
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024)
Monomial Matrix Group Equivariant Neural Functional Networks
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024)
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024)
Quasi-Equivariant Metanetworks
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2026)
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2026)
A Clifford Algebraic Approach to E(n)-Equivariant High-order Graph Neural Networks
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024)
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024)
Statistical Inference for Clustering-based Anomaly Detection
von: Phu, Nguyen Thi Minh, et al.
Veröffentlicht: (2025)
von: Phu, Nguyen Thi Minh, et al.
Veröffentlicht: (2025)
Improving Time Series Encoding with Noise-Aware Self-Supervised Learning and an Efficient Encoder
von: Nguyen, Duy A., et al.
Veröffentlicht: (2023)
von: Nguyen, Duy A., et al.
Veröffentlicht: (2023)
MMA: A Momentum Mamba Architecture for Human Activity Recognition with Inertial Sensors
von: Nguyen, Thai-Khanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thai-Khanh, et al.
Veröffentlicht: (2025)
Causal-Aware Generative Adversarial Networks with Reinforcement Learning
von: Nguyen, Tu Anh Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Tu Anh Hoang, et al.
Veröffentlicht: (2025)
Learning bridge numbers of knots
von: Vo, Hanh, et al.
Veröffentlicht: (2024)
von: Vo, Hanh, et al.
Veröffentlicht: (2024)
Defect Prediction with Content-based Features
von: Pham, Hung Viet, et al.
Veröffentlicht: (2024)
von: Pham, Hung Viet, et al.
Veröffentlicht: (2024)
MIST: Reliable Streaming Decision Trees for Online Class-Incremental Learning via McDiarmid Bound
von: Pham, Phu-Hoa, et al.
Veröffentlicht: (2026)
von: Pham, Phu-Hoa, et al.
Veröffentlicht: (2026)
Using Synthetic Data to estimate the True Error is theoretically and practically doable
von: Thanh, Hai Hoang, et al.
Veröffentlicht: (2025)
von: Thanh, Hai Hoang, et al.
Veröffentlicht: (2025)
Revisiting LARS for Large Batch Training Generalization of Neural Networks
von: Do, Khoi, et al.
Veröffentlicht: (2023)
von: Do, Khoi, et al.
Veröffentlicht: (2023)
Unlocking Compositional Generalization in Continual Few-Shot Learning
von: Nguyen-Lam, Phu-Quy, et al.
Veröffentlicht: (2026)
von: Nguyen-Lam, Phu-Quy, et al.
Veröffentlicht: (2026)
Spherical Tree-Sliced Wasserstein Distance
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2025)
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2025)
Cost-Adaptive Recourse Recommendation by Adaptive Preference Elicitation
von: Nguyen, Duy, et al.
Veröffentlicht: (2024)
von: Nguyen, Duy, et al.
Veröffentlicht: (2024)
More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning
von: Tran, Luong, et al.
Veröffentlicht: (2025)
von: Tran, Luong, et al.
Veröffentlicht: (2025)
Statistical Inference for Autoencoder-based Anomaly Detection after Representation Learning-based Domain Adaptation
von: Kiet, Tran Tuan, et al.
Veröffentlicht: (2025)
von: Kiet, Tran Tuan, et al.
Veröffentlicht: (2025)
Knowledge Abstraction for Knowledge-based Semantic Communication: A Generative Causality Invariant Approach
von: Nguyen, Minh-Duong, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh-Duong, et al.
Veröffentlicht: (2025)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
von: Pham, Loc, et al.
Veröffentlicht: (2025)
von: Pham, Loc, et al.
Veröffentlicht: (2025)
Revisiting Kernel Attention with Correlated Gaussian Process Representation
von: Bui, Long Minh, et al.
Veröffentlicht: (2025)
von: Bui, Long Minh, et al.
Veröffentlicht: (2025)
Coverage-Validity-Aware Algorithmic Recourse
von: Bui, Ngoc, et al.
Veröffentlicht: (2023)
von: Bui, Ngoc, et al.
Veröffentlicht: (2023)
Bayesian Optimization for Unknown Cost-Varying Variable Subsets with No-Regret Costs
von: Hoang, Vu Viet, et al.
Veröffentlicht: (2024)
von: Hoang, Vu Viet, et al.
Veröffentlicht: (2024)
Tree-Sliced Wasserstein Distance: A Geometric Perspective
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024)
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024)
Statistical Inference for Sequential Feature Selection after Domain Adaptation
von: Loc, Duong Tan, et al.
Veröffentlicht: (2025)
von: Loc, Duong Tan, et al.
Veröffentlicht: (2025)
Statistical Inference for Feature Selection after Optimal Transport-based Domain Adaptation
von: Loi, Nguyen Thang, et al.
Veröffentlicht: (2024)
von: Loi, Nguyen Thang, et al.
Veröffentlicht: (2024)
Understand the Effect of Importance Weighting in Deep Learning on Dataset Shift
von: Vo, Thien Nhan
Veröffentlicht: (2025)
von: Vo, Thien Nhan
Veröffentlicht: (2025)
LINKER: Learning Interactions Between Functional Groups and Residues With Chemical Knowledge-Enhanced Reasoning and Explainability
von: Pham, Phuc, et al.
Veröffentlicht: (2025)
von: Pham, Phuc, et al.
Veröffentlicht: (2025)
Vehicle Routing Problems via Quantum Graph Attention Network Deep Reinforcement Learning
von: Giang, Le Tung, et al.
Veröffentlicht: (2025)
von: Giang, Le Tung, et al.
Veröffentlicht: (2025)
Distributional Surgery for Language Model Activations
von: Nguyen, Bao, et al.
Veröffentlicht: (2025)
von: Nguyen, Bao, et al.
Veröffentlicht: (2025)
Range-aware Positional Encoding via High-order Pretraining: Theory and Practice
von: Nguyen, Viet Anh, et al.
Veröffentlicht: (2024)
von: Nguyen, Viet Anh, et al.
Veröffentlicht: (2024)
A Statistical Theory of Gated Attention through the Lens of Hierarchical Mixture of Experts
von: Nguyen, Viet, et al.
Veröffentlicht: (2026)
von: Nguyen, Viet, et al.
Veröffentlicht: (2026)
Statistical Inference for Clustering‐Based Anomaly Detection
von: Nguyen Thi Minh Phu, et al.
Veröffentlicht: (2025)
von: Nguyen Thi Minh Phu, et al.
Veröffentlicht: (2025)
Accelerating Transformers with Spectrum-Preserving Token Merging
von: Tran, Hoai-Chau, et al.
Veröffentlicht: (2024)
von: Tran, Hoai-Chau, et al.
Veröffentlicht: (2024)
Tree-Sliced Wasserstein Distance with Nonlinear Projection
von: Tran, Thanh, et al.
Veröffentlicht: (2025)
von: Tran, Thanh, et al.
Veröffentlicht: (2025)
CASUAL: Conditional Support Alignment for Domain Adaptation with Label Shift
von: Nguyen, Anh T, et al.
Veröffentlicht: (2023)
von: Nguyen, Anh T, et al.
Veröffentlicht: (2023)
Sparse Partial Optimal Transport via Quadratic Regularization
von: Tran, Khang, et al.
Veröffentlicht: (2025)
von: Tran, Khang, et al.
Veröffentlicht: (2025)
High Dimensional Bayesian Optimization using Lasso Variable Selection
von: Hoang, Vu Viet, et al.
Veröffentlicht: (2025)
von: Hoang, Vu Viet, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Demystifying the Token Dynamics of Deep Selective State Space Models
von: Vo, Thieu N, et al.
Veröffentlicht: (2024) -
Equivariant Polynomial Functional Networks
von: Vo, Thieu N., et al.
Veröffentlicht: (2024) -
Equivariant Neural Functional Networks for Transformers
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024) -
Monomial Matrix Group Equivariant Neural Functional Networks
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2024) -
Quasi-Equivariant Metanetworks
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2026)