Stabilizing Transformer Training Through Consensus
Fuente:
arXiv
Saved in:
| Main Authors: | Venkatasubramanian, Shyam, Moushegian, Sean, Lin, Michael, Park, Mir, Singhal, Ankit, Lee, Connor |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learn2Mix: Training Neural Networks Using Adaptive Data Integration
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
Diffusion-Based Hypothesis Testing and Change-Point Detection
by: Moushegian, Sean, et al.
Published: (2025)
by: Moushegian, Sean, et al.
Published: (2025)
Random Linear Projections Loss for Hyperplane-Based Optimization in Neural Networks
by: Venkatasubramanian, Shyam, et al.
Published: (2023)
by: Venkatasubramanian, Shyam, et al.
Published: (2023)
Steinmetz Neural Networks for Complex-Valued Data
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
Data-Driven Target Localization: Benchmarking Gradient Descent Using the Cramer-Rao Bound
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
Robust Score-Based Quickest Change Detection
by: Moushegian, Sean, et al.
Published: (2024)
by: Moushegian, Sean, et al.
Published: (2024)
RASPNet: A Benchmark Dataset for Radar Adaptive Signal Processing Applications
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
by: Venkatasubramanian, Shyam, et al.
Published: (2024)
Symmetry Breaking in Transformers for Efficient and Interpretable Training
by: Silverstein, Eva, et al.
Published: (2026)
by: Silverstein, Eva, et al.
Published: (2026)
To Pool or Not To Pool: Analyzing the Regularizing Effects of Group-Fair Training on Shared Models
by: Cousins, Cyrus, et al.
Published: (2024)
by: Cousins, Cyrus, et al.
Published: (2024)
Variational Encoder-Decoders for Learning Latent Representations of Physical Systems
by: Venkatasubramanian, Subashree, et al.
Published: (2024)
by: Venkatasubramanian, Subashree, et al.
Published: (2024)
PairNet: Training with Observed Pairs to Estimate Individual Treatment Effect
by: Nagalapatti, Lokesh, et al.
Published: (2024)
by: Nagalapatti, Lokesh, et al.
Published: (2024)
Learning a Consensus Sub-Network with Polarization Regularization and One Pass Training
by: Zhi, Xiaoying, et al.
Published: (2023)
by: Zhi, Xiaoying, et al.
Published: (2023)
Regularization-based Framework for Quantization-, Fault- and Variability-Aware Training
by: Biswas, Anmol, et al.
Published: (2025)
by: Biswas, Anmol, et al.
Published: (2025)
Leveraging Manifold Embeddings for Enhanced Graph Transformer Representations and Learning
by: Jyothish, Ankit, et al.
Published: (2025)
by: Jyothish, Ankit, et al.
Published: (2025)
CLOUD: A Scalable and Physics-Informed Foundation Model for Crystal Representation Learning
by: Xu, Changwen, et al.
Published: (2025)
by: Xu, Changwen, et al.
Published: (2025)
FieldFormer: Locality-Aware Transformers for Spatio-Temporal Modeling on Sparse Sensor Networks
by: Bhardwaj, Ankit, et al.
Published: (2025)
by: Bhardwaj, Ankit, et al.
Published: (2025)
Self-supervised Pretraining for Partial Differential Equations
by: Madhavan, Varun, et al.
Published: (2024)
by: Madhavan, Varun, et al.
Published: (2024)
Mean-Field Limits for Two-Layer Neural Networks Trained with Consensus-Based Optimization
by: De Deyn, William, et al.
Published: (2025)
by: De Deyn, William, et al.
Published: (2025)
Cascade-Aware Training of Language Models
by: Wang, Congchao, et al.
Published: (2024)
by: Wang, Congchao, et al.
Published: (2024)
Training Transformers for KV Cache Compressibility
by: Gelberg, Yoav, et al.
Published: (2026)
by: Gelberg, Yoav, et al.
Published: (2026)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
ConceptSearch: Towards Efficient Program Search Using LLMs for Abstraction and Reasoning Corpus (ARC)
by: Singhal, Kartik, et al.
Published: (2024)
by: Singhal, Kartik, et al.
Published: (2024)
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
by: Bylinkin, Dmitry, et al.
Published: (2025)
by: Bylinkin, Dmitry, et al.
Published: (2025)
Stability Selection via Variable Decorrelation
by: Nouraie, Mahdi, et al.
Published: (2025)
by: Nouraie, Mahdi, et al.
Published: (2025)
Breaking Symmetry When Training Transformers
by: Zuo, Chunsheng, et al.
Published: (2024)
by: Zuo, Chunsheng, et al.
Published: (2024)
Adversarial Training of Reward Models
by: Bukharin, Alexander, et al.
Published: (2025)
by: Bukharin, Alexander, et al.
Published: (2025)
Dataset-Driven Channel Masks in Transformers for Multivariate Time Series
by: Lee, Seunghan, et al.
Published: (2024)
by: Lee, Seunghan, et al.
Published: (2024)
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
by: Wang, Chao, et al.
Published: (2026)
by: Wang, Chao, et al.
Published: (2026)
Energy-Efficient Dynamic Training and Inference for GNN-Based Network Modeling
by: Singhal, Chetna, et al.
Published: (2025)
by: Singhal, Chetna, et al.
Published: (2025)
FLARE: One-Shot PE-Level Fault Localization in Systolic Arrays via Algebraic Test Vectors
by: Venkatasubramanian, Logashree, et al.
Published: (2026)
by: Venkatasubramanian, Logashree, et al.
Published: (2026)
Learning second-order TVD flux limiters using differentiable solvers
by: Huang, Chenyang, et al.
Published: (2025)
by: Huang, Chenyang, et al.
Published: (2025)
Impact of Layer Norm on Memorization and Generalization in Transformers
by: Singhal, Rishi, et al.
Published: (2025)
by: Singhal, Rishi, et al.
Published: (2025)
Topology-Informed Graph Transformer
by: Choi, Yun Young, et al.
Published: (2024)
by: Choi, Yun Young, et al.
Published: (2024)
On The Adaptation of Unlimiformer for Decoder-Only Transformers
by: Ahrabian, Kian, et al.
Published: (2024)
by: Ahrabian, Kian, et al.
Published: (2024)
FINDER: Stochastic Mirroring of Noisy Quasi-Newton Search and Deep Network Training
by: Suman, Uttam, et al.
Published: (2024)
by: Suman, Uttam, et al.
Published: (2024)
CViT: Continuous Vision Transformer for Operator Learning
by: Wang, Sifan, et al.
Published: (2024)
by: Wang, Sifan, et al.
Published: (2024)
Stability Analysis of Sharpness-Aware Minimization
by: Kim, Hoki, et al.
Published: (2023)
by: Kim, Hoki, et al.
Published: (2023)
Discriminative Ordering Through Ensemble Consensus
by: Ohl, Louis, et al.
Published: (2025)
by: Ohl, Louis, et al.
Published: (2025)
Stability and Generalization in Free Adversarial Training
by: Cheng, Xiwei, et al.
Published: (2024)
by: Cheng, Xiwei, et al.
Published: (2024)
Generalizable data-driven turbulence closure modeling on unstructured grids with differentiable physics
by: Kim, Hojin, et al.
Published: (2023)
by: Kim, Hojin, et al.
Published: (2023)
Similar Items
-
Learn2Mix: Training Neural Networks Using Adaptive Data Integration
by: Venkatasubramanian, Shyam, et al.
Published: (2024) -
Diffusion-Based Hypothesis Testing and Change-Point Detection
by: Moushegian, Sean, et al.
Published: (2025) -
Random Linear Projections Loss for Hyperplane-Based Optimization in Neural Networks
by: Venkatasubramanian, Shyam, et al.
Published: (2023) -
Steinmetz Neural Networks for Complex-Valued Data
by: Venkatasubramanian, Shyam, et al.
Published: (2024) -
Data-Driven Target Localization: Benchmarking Gradient Descent Using the Cramer-Rao Bound
by: Venkatasubramanian, Shyam, et al.
Published: (2024)