How Smooth Is Attention?
Fuente:
arXiv
Guardado en:
| Autores principales: | Castin, Valérie, Ablin, Pierre, Peyré, Gabriel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence
por: Castin, Valérie, et al.
Publicado: (2026)
por: Castin, Valérie, et al.
Publicado: (2026)
A Unified Perspective on the Dynamics of Deep Transformers
por: Castin, Valérie, et al.
Publicado: (2025)
por: Castin, Valérie, et al.
Publicado: (2025)
Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization
por: Ye, Zhenzhang, et al.
Publicado: (2024)
por: Ye, Zhenzhang, et al.
Publicado: (2024)
Robust Sublinear Convergence Rates for Iterative Bregman Projections
por: Peyré, Gabriel
Publicado: (2026)
por: Peyré, Gabriel
Publicado: (2026)
Muon Dynamics as a Spectral Wasserstein Flow
por: Peyré, Gabriel
Publicado: (2026)
por: Peyré, Gabriel
Publicado: (2026)
Optimal and Diffusion Transports in Machine Learning
por: Peyré, Gabriel
Publicado: (2025)
por: Peyré, Gabriel
Publicado: (2025)
Optimal Transport for Machine Learners
por: Peyré, Gabriel
Publicado: (2025)
por: Peyré, Gabriel
Publicado: (2025)
Token Sample Complexity of Attention
por: Bohbot, Léa, et al.
Publicado: (2025)
por: Bohbot, Léa, et al.
Publicado: (2025)
Nectar: Neural Estimation of Cached-Token Attention via Regression
por: Monteiro, João, et al.
Publicado: (2026)
por: Monteiro, João, et al.
Publicado: (2026)
The AdEMAMix Optimizer: Better, Faster, Older
por: Pagliardini, Matteo, et al.
Publicado: (2024)
por: Pagliardini, Matteo, et al.
Publicado: (2024)
Towards Understanding the Universality of Transformers for Next-Token Prediction
por: Sander, Michael E., et al.
Publicado: (2024)
por: Sander, Michael E., et al.
Publicado: (2024)
How do Transformers perform In-Context Autoregressive Learning?
por: Sander, Michael E., et al.
Publicado: (2024)
por: Sander, Michael E., et al.
Publicado: (2024)
Intrinsic training dynamics of deep neural networks
por: Marcotte, Sibylle, et al.
Publicado: (2025)
por: Marcotte, Sibylle, et al.
Publicado: (2025)
Geometry-Aware Discretization Error of Diffusion Models
por: Hurault, Samuel, et al.
Publicado: (2026)
por: Hurault, Samuel, et al.
Publicado: (2026)
Transformative or Conservative? Conservation laws for ResNets and Transformers
por: Marcotte, Sibylle, et al.
Publicado: (2025)
por: Marcotte, Sibylle, et al.
Publicado: (2025)
Dynamic Gradient Alignment for Online Data Mixing
por: Fan, Simin, et al.
Publicado: (2024)
por: Fan, Simin, et al.
Publicado: (2024)
How Smoothing is N-simplicial Attention?
por: Dussolle, Alexandre, et al.
Publicado: (2025)
por: Dussolle, Alexandre, et al.
Publicado: (2025)
Locking Pretrained Weights via Deep Low-Rank Residual Distillation
por: Sakamoto, Keitaro, et al.
Publicado: (2026)
por: Sakamoto, Keitaro, et al.
Publicado: (2026)
MVICAD2: Multi-View Independent Component Analysis with Delays and Dilations
por: Heurtebise, Ambroise, et al.
Publicado: (2025)
por: Heurtebise, Ambroise, et al.
Publicado: (2025)
Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
por: Marcotte, Sibylle, et al.
Publicado: (2023)
por: Marcotte, Sibylle, et al.
Publicado: (2023)
Learning from Samples: Inverse Problems over measures via Sharpened Fenchel-Young Losses
por: Andrade, Francisco, et al.
Publicado: (2025)
por: Andrade, Francisco, et al.
Publicado: (2025)
On the global convergence of gradient descent for wide shallow models with bounded nonlinearities
por: Petit, Romain, et al.
Publicado: (2026)
por: Petit, Romain, et al.
Publicado: (2026)
Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows
por: Marcotte, Sibylle, et al.
Publicado: (2024)
por: Marcotte, Sibylle, et al.
Publicado: (2024)
A Lower Bound and a Near-Optimal Algorithm for Bilevel Empirical Risk Minimization
por: Dagréou, Mathieu, et al.
Publicado: (2023)
por: Dagréou, Mathieu, et al.
Publicado: (2023)
Scaling Laws for Mixture Pretraining Under Data Constraints
por: Sedova, Anastasiia, et al.
Publicado: (2026)
por: Sedova, Anastasiia, et al.
Publicado: (2026)
A framework for bilevel optimization that enables stochastic and global variance reduction algorithms
por: Dagréou, Mathieu, et al.
Publicado: (2022)
por: Dagréou, Mathieu, et al.
Publicado: (2022)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
por: Grangier, David, et al.
Publicado: (2024)
por: Grangier, David, et al.
Publicado: (2024)
Need a Small Specialized Language Model? Plan Early!
por: Grangier, David, et al.
Publicado: (2024)
por: Grangier, David, et al.
Publicado: (2024)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
por: Ablin, Pierre, et al.
Publicado: (2025)
por: Ablin, Pierre, et al.
Publicado: (2025)
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
por: Barboni, Raphaël, et al.
Publicado: (2024)
por: Barboni, Raphaël, et al.
Publicado: (2024)
Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
por: Barboni, Raphaël, et al.
Publicado: (2025)
por: Barboni, Raphaël, et al.
Publicado: (2025)
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
por: Ramapuram, Jason, et al.
Publicado: (2024)
por: Ramapuram, Jason, et al.
Publicado: (2024)
Infeasible Deterministic, Stochastic, and Variance-Reduction Algorithms for Optimization under Orthogonality Constraints
por: Ablin, Pierre, et al.
Publicado: (2023)
por: Ablin, Pierre, et al.
Publicado: (2023)
Transformers are Universal In-context Learners
por: Furuya, Takashi, et al.
Publicado: (2024)
por: Furuya, Takashi, et al.
Publicado: (2024)
Playing Markov Games Without Observing Payoffs
por: Ablin, Daniel, et al.
Publicado: (2025)
por: Ablin, Daniel, et al.
Publicado: (2025)
DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
por: Gualdoni, Eleonora, et al.
Publicado: (2026)
por: Gualdoni, Eleonora, et al.
Publicado: (2026)
Careful with that Scalpel: Improving Gradient Surgery with an EMA
por: Hsieh, Yu-Guan, et al.
Publicado: (2024)
por: Hsieh, Yu-Guan, et al.
Publicado: (2024)
From Score Matching to Diffusion: A Fine-Grained Error Analysis in the Gaussian Setting
por: Hurault, Samuel, et al.
Publicado: (2025)
por: Hurault, Samuel, et al.
Publicado: (2025)
Learning Elastic Costs to Shape Monge Displacements
por: Klein, Michal, et al.
Publicado: (2023)
por: Klein, Michal, et al.
Publicado: (2023)
Scaling Categorical Flow Maps
por: Davis, Oscar, et al.
Publicado: (2026)
por: Davis, Oscar, et al.
Publicado: (2026)
Ejemplares similares
-
Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence
por: Castin, Valérie, et al.
Publicado: (2026) -
A Unified Perspective on the Dynamics of Deep Transformers
por: Castin, Valérie, et al.
Publicado: (2025) -
Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization
por: Ye, Zhenzhang, et al.
Publicado: (2024) -
Robust Sublinear Convergence Rates for Iterative Bregman Projections
por: Peyré, Gabriel
Publicado: (2026) -
Muon Dynamics as a Spectral Wasserstein Flow
por: Peyré, Gabriel
Publicado: (2026)