Faster Vision Mamba is Rebuilt in Minutes via Merged Token Re-training
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Mingjia, Zhou, Yuhao, Yu, Ruiji, Li, Zekai, Liang, Zhiyuan, Zhao, Xuanlei, Peng, Xiaojiang, Vedantam, Shanmukha Ramakrishna, Zhao, Wangbo, Wang, Kai, You, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Generalization Bound for a Family of Implicit Networks
by: Fung, Samy Wu, et al.
Published: (2024)
by: Fung, Samy Wu, et al.
Published: (2024)
Dynamic Vision Mamba
by: Wu, Mengxuan, et al.
Published: (2025)
by: Wu, Mengxuan, et al.
Published: (2025)
Saliency-Aware Multi-Route Thinking: Revisiting Vision-Language Reasoning
by: Shi, Mingjia, et al.
Published: (2026)
by: Shi, Mingjia, et al.
Published: (2026)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023)
by: Brännvall, Rickard, et al.
Published: (2023)
Can Sound Replace Vision in LLaVA With Token Substitution?
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
Mamba for Scalable and Efficient Personalized Recommendations
by: Starnes, Andrew, et al.
Published: (2024)
by: Starnes, Andrew, et al.
Published: (2024)
Adynamical systems view of training generativemodels and the memorization phenomenon
by: Athreya, Siva, et al.
Published: (2026)
by: Athreya, Siva, et al.
Published: (2026)
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
by: Yamchote, Phaphontee, et al.
Published: (2025)
by: Yamchote, Phaphontee, et al.
Published: (2025)
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization
by: Pavlov, Gorgi
Published: (2026)
by: Pavlov, Gorgi
Published: (2026)
Attention to Mamba: A Recipe for Cross-Architecture Distillation
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
MambaNetLK: Enhancing Colonoscopy Point Cloud Registration with Mamba
by: Jiang, Linzhe, et al.
Published: (2025)
by: Jiang, Linzhe, et al.
Published: (2025)
Differentiable Neural Networks with RePU Activation: with Applications to Score Estimation and Isotonic Regression
by: Shen, Guohao, et al.
Published: (2023)
by: Shen, Guohao, et al.
Published: (2023)
Attention Maps in 3D Shape Classification for Dental Stage Estimation with Class Node Graph Attention Networks
by: Buyukcakir, Barkin, et al.
Published: (2025)
by: Buyukcakir, Barkin, et al.
Published: (2025)
Marrying Compressed Sensing and Deep Signal Separation
by: Hickok, Truman, et al.
Published: (2024)
by: Hickok, Truman, et al.
Published: (2024)
Equidistribution-based training of Free Knot Splines and ReLU Neural Networks
by: Appella, Simone, et al.
Published: (2024)
by: Appella, Simone, et al.
Published: (2024)
Argus: Federated Non-convex Bilevel Learning over 6G Space-Air-Ground Integrated Network
by: Liu, Ya, et al.
Published: (2025)
by: Liu, Ya, et al.
Published: (2025)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
by: Nogales, Miguel, et al.
Published: (2025)
by: Nogales, Miguel, et al.
Published: (2025)
Temporal Attention Evolutional Graph Convolutional Network for Multivariate Time Series Forecasting
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
Cooperative Multi-Agent Deep Reinforcement Learning in Content Ranking Optimization
by: Qin, Zhou, et al.
Published: (2024)
by: Qin, Zhou, et al.
Published: (2024)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
by: Brusnicki, Roberto, et al.
Published: (2026)
by: Brusnicki, Roberto, et al.
Published: (2026)
Dual Preference Distribution Learning for Item Recommendation
by: Dong, Xue, et al.
Published: (2022)
by: Dong, Xue, et al.
Published: (2022)
Teaching and Learning under Deductive Errors
by: Telle, Jan Arne, et al.
Published: (2026)
by: Telle, Jan Arne, et al.
Published: (2026)
POA: Pre-training Once for Models of All Sizes
by: Zhang, Yingying, et al.
Published: (2024)
by: Zhang, Yingying, et al.
Published: (2024)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
by: Niizumi, Daisuke, et al.
Published: (2024)
by: Niizumi, Daisuke, et al.
Published: (2024)
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection
by: Niizumi, Daisuke, et al.
Published: (2024)
by: Niizumi, Daisuke, et al.
Published: (2024)
MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering
by: Luo, Yizhen, et al.
Published: (2024)
by: Luo, Yizhen, et al.
Published: (2024)
TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting
by: Stroh, Nicholas
Published: (2024)
by: Stroh, Nicholas
Published: (2024)
JVLGS: Joint Vision-Language Gas Leak Segmentation
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
HG-PIPE: Vision Transformer Acceleration with Hybrid-Grained Pipeline
by: Guo, Qingyu, et al.
Published: (2024)
by: Guo, Qingyu, et al.
Published: (2024)
Scalify: scale propagation for efficient low-precision LLM training
by: Balança, Paul, et al.
Published: (2024)
by: Balança, Paul, et al.
Published: (2024)
The Boundaries of Verifiable Accuracy, Robustness, and Generalisation in Deep Learning
by: Bastounis, Alexander, et al.
Published: (2023)
by: Bastounis, Alexander, et al.
Published: (2023)
OptiRoulette Optimizer: A New Stochastic Meta-Optimizer for up to 5.3x Faster Convergence
by: Mastromichalakis, Stamatis
Published: (2026)
by: Mastromichalakis, Stamatis
Published: (2026)
A Hybrid Deep Learning and Anomaly Detection Framework for Real-Time Malicious URL Classification
by: Khaled, Berkani, et al.
Published: (2025)
by: Khaled, Berkani, et al.
Published: (2025)
Benchmarking Vision Language Models on German Factual Data
by: Peinl, René, et al.
Published: (2025)
by: Peinl, René, et al.
Published: (2025)
Accuracy and stability of Artificial Neural Networks for HP-Splines frequency parameter selection
by: Bruni, Vittoria, et al.
Published: (2026)
by: Bruni, Vittoria, et al.
Published: (2026)
Deriving Decoder-Free Sparse Autoencoders from First Principles
by: Oursland, Alan
Published: (2026)
by: Oursland, Alan
Published: (2026)
DeepDRK: Deep Dependency Regularized Knockoff for Feature Selection
by: Shen, Hongyu, et al.
Published: (2024)
by: Shen, Hongyu, et al.
Published: (2024)
Directional Convergence, Benign Overfitting of Gradient Descent in leaky ReLU two-layer Neural Networks
by: Hashimoto, Ichiro
Published: (2025)
by: Hashimoto, Ichiro
Published: (2025)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
by: Guichoux, Téo, et al.
Published: (2025)
by: Guichoux, Téo, et al.
Published: (2025)
Similar Items
-
A Generalization Bound for a Family of Implicit Networks
by: Fung, Samy Wu, et al.
Published: (2024) -
Dynamic Vision Mamba
by: Wu, Mengxuan, et al.
Published: (2025) -
Saliency-Aware Multi-Route Thinking: Revisiting Vision-Language Reasoning
by: Shi, Mingjia, et al.
Published: (2026) -
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023) -
Can Sound Replace Vision in LLaVA With Token Substitution?
by: Vosoughi, Ali, et al.
Published: (2025)