Hyper-SET: Designing Transformers via Hyperspherical Energy Minimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Yunzhe, Zou, Difan, Xu, Dong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
di: Hu, Yunzhe, et al.
Pubblicazione: (2024)
di: Hu, Yunzhe, et al.
Pubblicazione: (2024)
Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKA
di: Smerkous, David, et al.
Pubblicazione: (2024)
di: Smerkous, David, et al.
Pubblicazione: (2024)
HyperCore: Coreset Selection under Noise via Hypersphere Models
di: Moser, Brian B., et al.
Pubblicazione: (2025)
di: Moser, Brian B., et al.
Pubblicazione: (2025)
What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
di: Chen, Xingwu, et al.
Pubblicazione: (2024)
di: Chen, Xingwu, et al.
Pubblicazione: (2024)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
di: Chen, Xingwu, et al.
Pubblicazione: (2024)
di: Chen, Xingwu, et al.
Pubblicazione: (2024)
Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference
di: Han, Yujin, et al.
Pubblicazione: (2024)
di: Han, Yujin, et al.
Pubblicazione: (2024)
Physics-Informed Neural PDE Solvers via Spatio-Temporal MeanFlow
di: Bai, Hanru, et al.
Pubblicazione: (2026)
di: Bai, Hanru, et al.
Pubblicazione: (2026)
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
Energy-Balanced Hyperspherical Graph Representation Learning via Structural Binding and Entropic Dispersion
di: Chen, Rui, et al.
Pubblicazione: (2025)
di: Chen, Rui, et al.
Pubblicazione: (2025)
Faster Sampling via Stochastic Gradient Proximal Sampler
di: Huang, Xunpeng, et al.
Pubblicazione: (2024)
di: Huang, Xunpeng, et al.
Pubblicazione: (2024)
Almost Linear Convergence under Minimal Score Assumptions: Quantized Transition Diffusion
di: Huang, Xunpeng, et al.
Pubblicazione: (2025)
di: Huang, Xunpeng, et al.
Pubblicazione: (2025)
Uncertainty Estimation via Hyperspherical Confidence Mapping
di: Choi, Eunseo, et al.
Pubblicazione: (2026)
di: Choi, Eunseo, et al.
Pubblicazione: (2026)
nGPT: Normalized Transformer with Representation Learning on the Hypersphere
di: Loshchilov, Ilya, et al.
Pubblicazione: (2024)
di: Loshchilov, Ilya, et al.
Pubblicazione: (2024)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
di: Jiang, Bingqing, et al.
Pubblicazione: (2026)
di: Jiang, Bingqing, et al.
Pubblicazione: (2026)
On the Memorization of Consistency Distillation for Diffusion Models
di: Jiang, Bingqing, et al.
Pubblicazione: (2026)
di: Jiang, Bingqing, et al.
Pubblicazione: (2026)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
di: Chen, Xingwu, et al.
Pubblicazione: (2025)
di: Chen, Xingwu, et al.
Pubblicazione: (2025)
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
Hierarchical Koopman Diffusion: Fast Generation with Interpretable Diffusion Trajectory
di: Bai, Hanru, et al.
Pubblicazione: (2025)
di: Bai, Hanru, et al.
Pubblicazione: (2025)
The Implicit Bias of Adam on Separable Data
di: Zhang, Chenyang, et al.
Pubblicazione: (2024)
di: Zhang, Chenyang, et al.
Pubblicazione: (2024)
Improving Implicit Regularization of SGD with Preconditioning for Least Square Problems
di: Su, Junwei, et al.
Pubblicazione: (2024)
di: Su, Junwei, et al.
Pubblicazione: (2024)
On the Limitation and Experience Replay for GNNs in Continual Learning
di: Su, Junwei, et al.
Pubblicazione: (2023)
di: Su, Junwei, et al.
Pubblicazione: (2023)
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
di: Li, Jichu, et al.
Pubblicazione: (2026)
di: Li, Jichu, et al.
Pubblicazione: (2026)
PRES: Toward Scalable Memory-Based Dynamic Graph Neural Networks
di: Su, Junwei, et al.
Pubblicazione: (2024)
di: Su, Junwei, et al.
Pubblicazione: (2024)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
Neural Collapse by Design: Learning Class Prototypes on the Hypersphere
di: Koromilas, Panagiotis, et al.
Pubblicazione: (2026)
di: Koromilas, Panagiotis, et al.
Pubblicazione: (2026)
Faster Sampling without Isoperimetry via Diffusion-based Monte Carlo
di: Huang, Xunpeng, et al.
Pubblicazione: (2024)
di: Huang, Xunpeng, et al.
Pubblicazione: (2024)
HYPO: Hyperspherical Out-of-Distribution Generalization
di: Bai, Haoyue, et al.
Pubblicazione: (2024)
di: Bai, Haoyue, et al.
Pubblicazione: (2024)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
Capturing Conditional Dependence via Auto-regressive Diffusion Models
di: Huang, Xunpeng, et al.
Pubblicazione: (2025)
di: Huang, Xunpeng, et al.
Pubblicazione: (2025)
Hyperspherical Normalization for Scalable Deep Reinforcement Learning
di: Lee, Hojoon, et al.
Pubblicazione: (2025)
di: Lee, Hojoon, et al.
Pubblicazione: (2025)
An Improved Analysis of Langevin Algorithms with Prior Diffusion for Non-Log-Concave Sampling
di: Huang, Xunpeng, et al.
Pubblicazione: (2024)
di: Huang, Xunpeng, et al.
Pubblicazione: (2024)
Learning under Quantization for High-Dimensional Linear Regression
di: Zhang, Dechen, et al.
Pubblicazione: (2025)
di: Zhang, Dechen, et al.
Pubblicazione: (2025)
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
di: Chen, Xingwu, et al.
Pubblicazione: (2025)
di: Chen, Xingwu, et al.
Pubblicazione: (2025)
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
di: Tang, Xuan, et al.
Pubblicazione: (2025)
di: Tang, Xuan, et al.
Pubblicazione: (2025)
On the Robustness of Transformers against Context Hijacking for Linear Classification
di: Li, Tianle, et al.
Pubblicazione: (2025)
di: Li, Tianle, et al.
Pubblicazione: (2025)
Hyperspherical Forward-Forward with Prototypical Representations
di: Sarode, Shalini, et al.
Pubblicazione: (2026)
di: Sarode, Shalini, et al.
Pubblicazione: (2026)
F-Adapter: Frequency-Adaptive Parameter-Efficient Fine-Tuning in Scientific Machine Learning
di: Zhang, Hangwei, et al.
Pubblicazione: (2025)
di: Zhang, Hangwei, et al.
Pubblicazione: (2025)
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks
di: Zhang, Chenyang, et al.
Pubblicazione: (2025)
di: Zhang, Chenyang, et al.
Pubblicazione: (2025)
Towards Robust Graph Incremental Learning on Evolving Graphs
di: Su, Junwei, et al.
Pubblicazione: (2024)
di: Su, Junwei, et al.
Pubblicazione: (2024)
On the Benefits of Over-parameterization for Out-of-Distribution Generalization
di: Hao, Yifan, et al.
Pubblicazione: (2024)
di: Hao, Yifan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
di: Hu, Yunzhe, et al.
Pubblicazione: (2024) -
Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKA
di: Smerkous, David, et al.
Pubblicazione: (2024) -
HyperCore: Coreset Selection under Noise via Hypersphere Models
di: Moser, Brian B., et al.
Pubblicazione: (2025) -
What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
di: Chen, Xingwu, et al.
Pubblicazione: (2024) -
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
di: Chen, Xingwu, et al.
Pubblicazione: (2024)