On the Training Convergence of Transformers for In-Context Classification of Gaussian Mixtures
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Wei, Zhou, Ruida, Yang, Jing, Shen, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cost-Aware Optimal Pairwise Pure Exploration
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Harnessing the Power of Federated Learning in Federated Contextual Bandits
by: Shi, Chengshuai, et al.
Published: (2023)
by: Shi, Chengshuai, et al.
Published: (2023)
Decision Feedback In-Context Learning for Wireless Symbol Detection
by: Fan, Li, et al.
Published: (2025)
by: Fan, Li, et al.
Published: (2025)
An Information-Theoretic Approach to Understanding Transformers' In-Context Learning of Variable-Order Markov Chains
by: Zhou, Ruida, et al.
Published: (2024)
by: Zhou, Ruida, et al.
Published: (2024)
Decision Feedback In-Context Symbol Detection over Block-Fading Channels
by: Fan, Li, et al.
Published: (2024)
by: Fan, Li, et al.
Published: (2024)
On the Convergence Analysis of Muon
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
Chain-of-Thought Enhanced Shallow Transformers for Wireless Symbol Detection
by: Fan, Li, et al.
Published: (2025)
by: Fan, Li, et al.
Published: (2025)
On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery
by: Liu, Renpu, et al.
Published: (2024)
by: Liu, Renpu, et al.
Published: (2024)
Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
Greedy Sampling Is Provably Efficient for RLHF
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
by: Wu, Di, et al.
Published: (2026)
by: Wu, Di, et al.
Published: (2026)
Expectation Maximization (EM) Converges for General Agnostic Mixtures
by: Ghosh, Avishek
Published: (2026)
by: Ghosh, Avishek
Published: (2026)
Stochastic Smoothed Gradient Descent Ascent for Federated Minimax Optimization
by: Shen, Wei, et al.
Published: (2023)
by: Shen, Wei, et al.
Published: (2023)
A Single-Loop First-Order Algorithm for Linearly Constrained Bilevel Optimization
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
Average Reward Reinforcement Learning for Wireless Radio Resource Management
by: Yang, Kun, et al.
Published: (2025)
by: Yang, Kun, et al.
Published: (2025)
Information Theoretic Learning for Diffusion Models with Warm Start
by: Shen, Yirong, et al.
Published: (2025)
by: Shen, Yirong, et al.
Published: (2025)
Variational Bayesian Methods for a Tree-Structured Stick-Breaking Process Mixture of Gaussians by Application of the Bayes Codes for Context Tree Models
by: Nakahara, Yuta
Published: (2024)
by: Nakahara, Yuta
Published: (2024)
Robust Federated Personalised Mean Estimation for the Gaussian Mixture Model
by: Managoli, Malhar A., et al.
Published: (2025)
by: Managoli, Malhar A., et al.
Published: (2025)
MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
by: Yang, Tong, et al.
Published: (2024)
by: Yang, Tong, et al.
Published: (2024)
Generalization Guarantees for Representation Learning via Data-Dependent Gaussian Mixture Priors
by: Sefidgaran, Milad, et al.
Published: (2025)
by: Sefidgaran, Milad, et al.
Published: (2025)
Confirmation Bias in Gaussian Mixture Models
by: Balanov, Amnon, et al.
Published: (2024)
by: Balanov, Amnon, et al.
Published: (2024)
Personalized Federated Learning with Attention-based Client Selection
by: Chen, Zihan, et al.
Published: (2023)
by: Chen, Zihan, et al.
Published: (2023)
Generalization Guarantees for Multi-View Representation Learning and Application to Regularization via Gaussian Product Mixture Prior
by: Sefidgaran, Milad, et al.
Published: (2025)
by: Sefidgaran, Milad, et al.
Published: (2025)
On optimal solutions of classical and sliced Wasserstein GANs with non-Gaussian data
by: Huang, Yu-Jui, et al.
Published: (2025)
by: Huang, Yu-Jui, et al.
Published: (2025)
An Autoencoder-Based Constellation Design for AirComp in Wireless Federated Learning
by: Mu, Yujia, et al.
Published: (2024)
by: Mu, Yujia, et al.
Published: (2024)
Data-adaptive Differentially Private Prompt Synthesis for In-Context Learning
by: Gao, Fengyu, et al.
Published: (2024)
by: Gao, Fengyu, et al.
Published: (2024)
A Convergence Analysis of Approximate Message Passing with Non-Separable Functions and Applications to Multi-Class Classification
by: Çakmak, Burak, et al.
Published: (2024)
by: Çakmak, Burak, et al.
Published: (2024)
Federated Split Learning with Improved Communication and Storage Efficiency
by: Mu, Yujia, et al.
Published: (2025)
by: Mu, Yujia, et al.
Published: (2025)
Constrained Gaussian Wasserstein Optimal Transport with Commutative Covariance Matrices
by: Chen, Jun, et al.
Published: (2025)
by: Chen, Jun, et al.
Published: (2025)
IRS-Enhanced Secure Semantic Communication Networks: Cross-Layer and Context-Awared Resource Allocation
by: Wang, Lingyi, et al.
Published: (2024)
by: Wang, Lingyi, et al.
Published: (2024)
MIMO Detection via Gaussian Mixture Expectation Propagation: A Bayesian Machine Learning Approach for High-Order High-Dimensional MIMO Systems
by: Shayovitz, Shachar, et al.
Published: (2024)
by: Shayovitz, Shachar, et al.
Published: (2024)
PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources
by: Park, Bumsu, et al.
Published: (2026)
by: Park, Bumsu, et al.
Published: (2026)
Random pairing MLE for estimation of item parameters in Rasch model
by: Yang, Yuepeng, et al.
Published: (2024)
by: Yang, Yuepeng, et al.
Published: (2024)
Sample efficient inductive matrix completion with noise and inexact side information
by: Yang, Yuepeng, et al.
Published: (2026)
by: Yang, Yuepeng, et al.
Published: (2026)
Towards Faster Non-Asymptotic Convergence for Diffusion-Based Generative Models
by: Li, Gen, et al.
Published: (2023)
by: Li, Gen, et al.
Published: (2023)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
by: Fesharaki, Amirmehdi Jafari, et al.
Published: (2026)
by: Fesharaki, Amirmehdi Jafari, et al.
Published: (2026)
On the best approximation by finite Gaussian mixtures
by: Ma, Yun, et al.
Published: (2024)
by: Ma, Yun, et al.
Published: (2024)
A Mixture of Experts Vision Transformer for High-Fidelity Surface Code Decoding
by: Nguyen, Hoang Viet, et al.
Published: (2026)
by: Nguyen, Hoang Viet, et al.
Published: (2026)
Digital Over-the-Air Federated Learning in Multi-Antenna Systems
by: Wang, Sihua, et al.
Published: (2023)
by: Wang, Sihua, et al.
Published: (2023)
Similar Items
-
Cost-Aware Optimal Pairwise Pure Exploration
by: Wu, Di, et al.
Published: (2025) -
Harnessing the Power of Federated Learning in Federated Contextual Bandits
by: Shi, Chengshuai, et al.
Published: (2023) -
Decision Feedback In-Context Learning for Wireless Symbol Detection
by: Fan, Li, et al.
Published: (2025) -
An Information-Theoretic Approach to Understanding Transformers' In-Context Learning of Variable-Order Markov Chains
by: Zhou, Ruida, et al.
Published: (2024) -
Decision Feedback In-Context Symbol Detection over Block-Fading Channels
by: Fan, Li, et al.
Published: (2024)