Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongkang, Lu, Songtao, Cui, Xiaodong, Chen, Pin-Yu, Wang, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
by: Li, Hongkang, et al.
Published: (2025)
by: Li, Hongkang, et al.
Published: (2025)
Theoretical Learning Performance of Graph Neural Networks: The Impact of Jumping Connections and Layer-wise Sparsification
by: Sun, Jiawei, et al.
Published: (2025)
by: Sun, Jiawei, et al.
Published: (2025)
How does promoting the minority fraction affect generalization? A theoretical study of the one-hidden-layer neural network on group imbalance
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning
by: Zhang, Shuai, et al.
Published: (2024)
by: Zhang, Shuai, et al.
Published: (2024)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
by: Park, Jongho, et al.
Published: (2024)
by: Park, Jongho, et al.
Published: (2024)
Mamba Modulation: On the Length Generalization of Mamba
by: Lu, Peng, et al.
Published: (2025)
by: Lu, Peng, et al.
Published: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
Is Mamba Capable of In-Context Learning?
by: Grazzi, Riccardo, et al.
Published: (2024)
by: Grazzi, Riccardo, et al.
Published: (2024)
Visual prompting reimagined: The power of the Activation Prompts
by: Zhang, Yihua, et al.
Published: (2026)
by: Zhang, Yihua, et al.
Published: (2026)
An Information-Theoretic Analysis of In-Context Learning
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs
by: Li, Hongkang, et al.
Published: (2026)
by: Li, Hongkang, et al.
Published: (2026)
Computational Safety for Generative AI: A Signal Processing Perspective
by: Chen, Pin-Yu
Published: (2025)
by: Chen, Pin-Yu
Published: (2025)
Towards the Effect of Examples on In-Context Learning: A Theoretical Case Study
by: He, Pengfei, et al.
Published: (2024)
by: He, Pengfei, et al.
Published: (2024)
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
by: Saif, A F M, et al.
Published: (2024)
by: Saif, A F M, et al.
Published: (2024)
Efficient Generation of Hidden Outliers for Improved Outlier Detection
by: Cribeiro-Ramallo, Jose, et al.
Published: (2024)
by: Cribeiro-Ramallo, Jose, et al.
Published: (2024)
Adversarial Imitation Learning with General Function Approximation: Theoretical Analysis and Practical Algorithms
by: Xu, Tian, et al.
Published: (2026)
by: Xu, Tian, et al.
Published: (2026)
Understanding Generalization and Forgetting in In-Context Continual Learning
by: Li, Guangyu, et al.
Published: (2026)
by: Li, Guangyu, et al.
Published: (2026)
Can Dense Connectivity Benefit Outlier Detection? An Odyssey with NAS
by: Fu, Hao, et al.
Published: (2024)
by: Fu, Hao, et al.
Published: (2024)
Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
by: Zhang, Xingxuan, et al.
Published: (2025)
by: Zhang, Xingxuan, et al.
Published: (2025)
Learning under Commission and Omission Event Outliers
by: Zhang, Yuecheng, et al.
Published: (2025)
by: Zhang, Yuecheng, et al.
Published: (2025)
Outlier Smoothing with Closed-Form Rotations for W4A4 Large Language Model Quantization
by: Xiao, Jinying, et al.
Published: (2025)
by: Xiao, Jinying, et al.
Published: (2025)
Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
by: Saif, A F M, et al.
Published: (2025)
by: Saif, A F M, et al.
Published: (2025)
Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured Data
by: Li, Binghui, et al.
Published: (2024)
by: Li, Binghui, et al.
Published: (2024)
Mamba-PTQ: Outlier Channels in Recurrent Large Language Models
by: Pierro, Alessandro, et al.
Published: (2024)
by: Pierro, Alessandro, et al.
Published: (2024)
A Single-Loop Gradient Descent and Perturbed Ascent Algorithm for Nonconvex Functional Constrained Optimization
by: Lu, Songtao
Published: (2022)
by: Lu, Songtao
Published: (2022)
Towards a Theoretical Understanding to the Generalization of RLHF
by: Li, Zhaochun, et al.
Published: (2026)
by: Li, Zhaochun, et al.
Published: (2026)
Optimization of Retrieval-Augmented Generation Context with Outlier Detection
by: Bulgakov, Vitaly
Published: (2024)
by: Bulgakov, Vitaly
Published: (2024)
Can DPO Learn Diverse Human Values? A Theoretical Scaling Law
by: Im, Shawn, et al.
Published: (2024)
by: Im, Shawn, et al.
Published: (2024)
Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
by: Yan, Tingkai, et al.
Published: (2025)
by: Yan, Tingkai, et al.
Published: (2025)
OWLed: Outlier-weighed Layerwise Pruning for Efficient Autonomous Driving Framework
by: Li, Jiaxi, et al.
Published: (2024)
by: Li, Jiaxi, et al.
Published: (2024)
Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
by: Chhabra, Anshuman, et al.
Published: (2024)
by: Chhabra, Anshuman, et al.
Published: (2024)
Thompson Sampling in Online RLHF with General Function Approximation
by: Feng, Songtao, et al.
Published: (2025)
by: Feng, Songtao, et al.
Published: (2025)
Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement Learning
by: Li, Lanqing, et al.
Published: (2024)
by: Li, Lanqing, et al.
Published: (2024)
Can We Theoretically Quantify the Impacts of Local Updates on the Generalization Performance of Federated Learning?
by: Ju, Peizhong, et al.
Published: (2024)
by: Ju, Peizhong, et al.
Published: (2024)
Similar Items
-
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
by: Li, Hongkang, et al.
Published: (2024) -
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024) -
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models
by: Shandirasegaran, Mugunthan, et al.
Published: (2026) -
What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
by: Li, Hongkang, et al.
Published: (2024) -
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
by: Li, Hongkang, et al.
Published: (2024)