DeepCrossAttention: Supercharging Transformer Residual Connections
Fuente:
arXiv
Guardado en:
| Autores principales: | Heddes, Mike, Javanmard, Adel, Axiotis, Kyriakos, Fu, Gang, Bateni, MohammadHossein, Mirrokni, Vahab |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SequentialAttention++ for Block Sparsification: Differentiable Pruning Meets Combinatorial Optimization
por: Yasuda, Taisuke, et al.
Publicado: (2024)
por: Yasuda, Taisuke, et al.
Publicado: (2024)
Sequential Attention for Feature Selection
por: Yasuda, Taisuke, et al.
Publicado: (2022)
por: Yasuda, Taisuke, et al.
Publicado: (2022)
Understanding the Role of Training Data in Test-Time Scaling
por: Javanmard, Adel, et al.
Publicado: (2025)
por: Javanmard, Adel, et al.
Publicado: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
por: Javanmard, Adel, et al.
Publicado: (2026)
por: Javanmard, Adel, et al.
Publicado: (2026)
PriorBoost: An Adaptive Algorithm for Learning from Aggregate Responses
por: Javanmard, Adel, et al.
Publicado: (2024)
por: Javanmard, Adel, et al.
Publicado: (2024)
Improving the Variance of Differentially Private Randomized Experiments through Clustering
por: Javanmard, Adel, et al.
Publicado: (2023)
por: Javanmard, Adel, et al.
Publicado: (2023)
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
por: Javanmard, Adel, et al.
Publicado: (2024)
por: Javanmard, Adel, et al.
Publicado: (2024)
Self-Boost via Optimal Retraining: An Analysis via Approximate Message Passing
por: Javanmard, Adel, et al.
Publicado: (2025)
por: Javanmard, Adel, et al.
Publicado: (2025)
Learning Rate Schedules in the Presence of Distribution Shift
por: Fahrbach, Matthew, et al.
Publicado: (2023)
por: Fahrbach, Matthew, et al.
Publicado: (2023)
Optimistic Rates for Learning from Label Proportions
por: Li, Gene, et al.
Publicado: (2024)
por: Li, Gene, et al.
Publicado: (2024)
Budget Allocation for Unknown Value Functions in a Lipschitz Space
por: Bateni, MohammadHossein, et al.
Publicado: (2025)
por: Bateni, MohammadHossein, et al.
Publicado: (2025)
Data-Efficient Learning via Clustering-Based Sensitivity Sampling: Foundation Models and Beyond
por: Axiotis, Kyriakos, et al.
Publicado: (2024)
por: Axiotis, Kyriakos, et al.
Publicado: (2024)
Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
por: Naharas, Nilay, et al.
Publicado: (2025)
por: Naharas, Nilay, et al.
Publicado: (2025)
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
por: Behrouz, Ali, et al.
Publicado: (2025)
por: Behrouz, Ali, et al.
Publicado: (2025)
Replicable Composition
por: Banihashem, Kiarash, et al.
Publicado: (2026)
por: Banihashem, Kiarash, et al.
Publicado: (2026)
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
por: Kacham, Praneeth, et al.
Publicado: (2023)
por: Kacham, Praneeth, et al.
Publicado: (2023)
Always-Sparse Training by Growing Connections with Guided Stochastic Exploration
por: Heddes, Mike, et al.
Publicado: (2024)
por: Heddes, Mike, et al.
Publicado: (2024)
Retraining with Predicted Hard Labels Provably Increases Model Accuracy
por: Das, Rudrajit, et al.
Publicado: (2024)
por: Das, Rudrajit, et al.
Publicado: (2024)
Approximately Optimal Core Shapes for Tensor Decompositions
por: Ghadiri, Mehrdad, et al.
Publicado: (2023)
por: Ghadiri, Mehrdad, et al.
Publicado: (2023)
A Scalable Algorithm for Individually Fair K-means Clustering
por: Bateni, MohammadHossein, et al.
Publicado: (2024)
por: Bateni, MohammadHossein, et al.
Publicado: (2024)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
por: Nguyen, Dang, et al.
Publicado: (2025)
por: Nguyen, Dang, et al.
Publicado: (2025)
Differentially Private Model-X Knockoffs via Johnson-Lindenstrauss Transform
por: Tao, Yuxuan, et al.
Publicado: (2025)
por: Tao, Yuxuan, et al.
Publicado: (2025)
Trellis: Learning to Compress Key-Value Memory in Attention Models
por: Karami, Mahdi, et al.
Publicado: (2025)
por: Karami, Mahdi, et al.
Publicado: (2025)
Differentially Private Synthetic Data Release for Topics API Outputs
por: Dick, Travis, et al.
Publicado: (2025)
por: Dick, Travis, et al.
Publicado: (2025)
Nested Learning: The Illusion of Deep Learning Architectures
por: Behrouz, Ali, et al.
Publicado: (2025)
por: Behrouz, Ali, et al.
Publicado: (2025)
Smooth Anonymity for Sparse Graphs
por: Epasto, Alessandro, et al.
Publicado: (2022)
por: Epasto, Alessandro, et al.
Publicado: (2022)
Networked Information Aggregation for Binary Classification
por: Bateni, MohammadHossein, et al.
Publicado: (2026)
por: Bateni, MohammadHossein, et al.
Publicado: (2026)
Optimal Approximation -- Smoothness Tradeoffs for Soft-Max Functions
por: Epasto, Alessandro, et al.
Publicado: (2020)
por: Epasto, Alessandro, et al.
Publicado: (2020)
Lattice: Learning to Efficiently Compress the Memory
por: Karami, Mahdi, et al.
Publicado: (2025)
por: Karami, Mahdi, et al.
Publicado: (2025)
SYNAPSE-G: Bridging Large Language Models and Graph Learning for Rare Event Classification
por: Tavakkol, Sasan, et al.
Publicado: (2025)
por: Tavakkol, Sasan, et al.
Publicado: (2025)
Replicable Clustering
por: Esfandiari, Hossein, et al.
Publicado: (2023)
por: Esfandiari, Hossein, et al.
Publicado: (2023)
The curse of overparametrization in adversarial training: Precise analysis of robust generalization for random features regression
por: Hassani, Hamed, et al.
Publicado: (2022)
por: Hassani, Hamed, et al.
Publicado: (2022)
PolarQuant: Quantizing KV Caches with Polar Transformation
por: Han, Insu, et al.
Publicado: (2025)
por: Han, Insu, et al.
Publicado: (2025)
Titans: Learning to Memorize at Test Time
por: Behrouz, Ali, et al.
Publicado: (2024)
por: Behrouz, Ali, et al.
Publicado: (2024)
Sampling and Loss Weights in Multi-Domain Training
por: Salmani, Mahdi, et al.
Publicado: (2025)
por: Salmani, Mahdi, et al.
Publicado: (2025)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
por: Das, Rudrajit, et al.
Publicado: (2026)
por: Das, Rudrajit, et al.
Publicado: (2026)
Multi-Task Dynamic Pricing in Credit Market with Contextual Information
por: Javanmard, Adel, et al.
Publicado: (2024)
por: Javanmard, Adel, et al.
Publicado: (2024)
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
por: Li, Zeman, et al.
Publicado: (2025)
por: Li, Zeman, et al.
Publicado: (2025)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
por: Karami, Mahdi, et al.
Publicado: (2025)
por: Karami, Mahdi, et al.
Publicado: (2025)
Supercharging Graph Transformers with Advective Diffusion
por: Wu, Qitian, et al.
Publicado: (2023)
por: Wu, Qitian, et al.
Publicado: (2023)
Ejemplares similares
-
SequentialAttention++ for Block Sparsification: Differentiable Pruning Meets Combinatorial Optimization
por: Yasuda, Taisuke, et al.
Publicado: (2024) -
Sequential Attention for Feature Selection
por: Yasuda, Taisuke, et al.
Publicado: (2022) -
Understanding the Role of Training Data in Test-Time Scaling
por: Javanmard, Adel, et al.
Publicado: (2025) -
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
por: Javanmard, Adel, et al.
Publicado: (2026) -
PriorBoost: An Adaptive Algorithm for Learning from Aggregate Responses
por: Javanmard, Adel, et al.
Publicado: (2024)