Swimba: Switch Mamba Model Scales State Space Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Du, Zhixu, Chitty-Venkata, Krishna Teja, Emani, Murali, Vishwanath, Venkatram, Li, Hai Helen, Chen, Yiran |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
por: Thapa, Krishu K, et al.
Publicado: (2025)
por: Thapa, Krishu K, et al.
Publicado: (2025)
BaKlaVa -- Budgeted Allocation of KV cache for Long-context Inference
por: Gulhan, Ahmed Burak, et al.
Publicado: (2025)
por: Gulhan, Ahmed Burak, et al.
Publicado: (2025)
PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
MoPEQ: Mixture of Mixed Precision Quantized Experts
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025)
LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2024)
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2024)
MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
por: Ma, Haiyue, et al.
Publicado: (2025)
por: Ma, Haiyue, et al.
Publicado: (2025)
Quality Measures for Dynamic Graph Generative Models
por: Hosseini, Ryien, et al.
Publicado: (2025)
por: Hosseini, Ryien, et al.
Publicado: (2025)
A Deep Probabilistic Framework for Continuous Time Dynamic Graph Generation
por: Hosseini, Ryien, et al.
Publicado: (2024)
por: Hosseini, Ryien, et al.
Publicado: (2024)
SiDA-MoE: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
por: Du, Zhixu, et al.
Publicado: (2023)
por: Du, Zhixu, et al.
Publicado: (2023)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
por: Shao, Zishan, et al.
Publicado: (2025)
por: Shao, Zishan, et al.
Publicado: (2025)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
por: Zhan, Zheng, et al.
Publicado: (2025)
por: Zhan, Zheng, et al.
Publicado: (2025)
AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
por: Hatanpää, Väinö, et al.
Publicado: (2025)
por: Hatanpää, Väinö, et al.
Publicado: (2025)
Sketch-Augmented Features Improve Learning Long-Range Dependencies in Graph Neural Networks
por: Hosseini, Ryien, et al.
Publicado: (2025)
por: Hosseini, Ryien, et al.
Publicado: (2025)
Mamba-3: Improved Sequence Modeling using State Space Principles
por: Lahoti, Aakash, et al.
Publicado: (2026)
por: Lahoti, Aakash, et al.
Publicado: (2026)
Extending $μ$P: Spectral Conditions for Feature Learning Across Optimizers
por: Gupta, Akshita, et al.
Publicado: (2026)
por: Gupta, Akshita, et al.
Publicado: (2026)
DyG-Mamba: Continuous State Space Modeling on Dynamic Graphs
por: Li, Dongyuan, et al.
Publicado: (2024)
por: Li, Dongyuan, et al.
Publicado: (2024)
Graph Mamba: Towards Learning on Graphs with State Space Models
por: Behrouz, Ali, et al.
Publicado: (2024)
por: Behrouz, Ali, et al.
Publicado: (2024)
Mamba State-Space Models Are Lyapunov-Stable Learners
por: Halloran, John T., et al.
Publicado: (2024)
por: Halloran, John T., et al.
Publicado: (2024)
PerfMamba: Performance Analysis and Pruning of Selective State Space Models
por: Asif, Abdullah Al, et al.
Publicado: (2025)
por: Asif, Abdullah Al, et al.
Publicado: (2025)
BarcodeMamba+: Advancing State-Space Models for Fungal Biodiversity Research
por: Gao, Tiancheng, et al.
Publicado: (2025)
por: Gao, Tiancheng, et al.
Publicado: (2025)
BarcodeMamba: State Space Models for Biodiversity Analysis
por: Gao, Tiancheng, et al.
Publicado: (2024)
por: Gao, Tiancheng, et al.
Publicado: (2024)
MemMamba: Rethinking Memory Patterns in State Space Model
por: Wang, Youjin, et al.
Publicado: (2025)
por: Wang, Youjin, et al.
Publicado: (2025)
MambaByte: Token-free Selective State Space Model
por: Wang, Junxiong, et al.
Publicado: (2024)
por: Wang, Junxiong, et al.
Publicado: (2024)
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
por: Gu, Albert, et al.
Publicado: (2023)
por: Gu, Albert, et al.
Publicado: (2023)
MambaLRP: Explaining Selective State Space Sequence Models
por: Jafari, Farnoush Rezaei, et al.
Publicado: (2024)
por: Jafari, Farnoush Rezaei, et al.
Publicado: (2024)
ss-Mamba: Semantic-Spline Selective State-Space Model
por: Ye, Zuochen
Publicado: (2025)
por: Ye, Zuochen
Publicado: (2025)
Foundation Models for Discovery and Exploration in Chemical Space
por: Wadell, Alexius, et al.
Publicado: (2025)
por: Wadell, Alexius, et al.
Publicado: (2025)
Deep Switching State Space Model (DS$^3$M) for Nonlinear Time Series Forecasting with Regime Switching
por: Xu, Xiuqin, et al.
Publicado: (2021)
por: Xu, Xiuqin, et al.
Publicado: (2021)
DG-Mamba: Robust and Efficient Dynamic Graph Structure Learning with Selective State Space Models
por: Yuan, Haonan, et al.
Publicado: (2024)
por: Yuan, Haonan, et al.
Publicado: (2024)
Decision Mamba: A Multi-Grained State Space Model with Self-Evolution Regularization for Offline RL
por: Lv, Qi, et al.
Publicado: (2024)
por: Lv, Qi, et al.
Publicado: (2024)
STG-Mamba: Spatial-Temporal Graph Learning via Selective State Space Model
por: Li, Lincan, et al.
Publicado: (2024)
por: Li, Lincan, et al.
Publicado: (2024)
FR-Mamba: Time-Series Physical Field Reconstruction Based on State Space Model
por: Long, Jiahuan, et al.
Publicado: (2025)
por: Long, Jiahuan, et al.
Publicado: (2025)
KalMamba: Towards Efficient Probabilistic State Space Models for RL under Uncertainty
por: Becker, Philipp, et al.
Publicado: (2024)
por: Becker, Philipp, et al.
Publicado: (2024)
SambaMixer: State of Health Prediction of Li-ion Batteries using Mamba State Space Models
por: Olalde-Verano, José Ignacio, et al.
Publicado: (2024)
por: Olalde-Verano, José Ignacio, et al.
Publicado: (2024)
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models
por: Shandirasegaran, Mugunthan, et al.
Publicado: (2026)
por: Shandirasegaran, Mugunthan, et al.
Publicado: (2026)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
por: Chen, Yifang, et al.
Publicado: (2024)
por: Chen, Yifang, et al.
Publicado: (2024)
Ejemplares similares
-
LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025) -
ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025) -
LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
por: Chitty-Venkata, Krishna Teja, et al.
Publicado: (2025) -
PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
por: Thapa, Krishu K, et al.
Publicado: (2025) -
BaKlaVa -- Budgeted Allocation of KV cache for Long-context Inference
por: Gulhan, Ahmed Burak, et al.
Publicado: (2025)