Famba-V: Fast Vision Mamba with Cross-Layer Token Fusion
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Hui, Wan, Zhongwei, Wang, Xin, Zhang, Mi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation
di: Liu, Che, et al.
Pubblicazione: (2024)
di: Liu, Che, et al.
Pubblicazione: (2024)
Fusion-Mamba for Cross-modality Object Detection
di: Dong, Wenhao, et al.
Pubblicazione: (2024)
di: Dong, Wenhao, et al.
Pubblicazione: (2024)
EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion
di: Yang, Zichuan, et al.
Pubblicazione: (2025)
di: Yang, Zichuan, et al.
Pubblicazione: (2025)
COMO: Cross-Mamba Interaction and Offset-Guided Fusion for Multimodal Object Detection
di: Liu, Chang, et al.
Pubblicazione: (2024)
di: Liu, Chang, et al.
Pubblicazione: (2024)
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
di: Qian, Chen, et al.
Pubblicazione: (2026)
di: Qian, Chen, et al.
Pubblicazione: (2026)
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
di: Zhu, Xuanyu, et al.
Pubblicazione: (2026)
di: Zhu, Xuanyu, et al.
Pubblicazione: (2026)
Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias
di: Wan, Zhongwei, et al.
Pubblicazione: (2023)
di: Wan, Zhongwei, et al.
Pubblicazione: (2023)
LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks
di: Liu, Hui, et al.
Pubblicazione: (2025)
di: Liu, Hui, et al.
Pubblicazione: (2025)
DepthMamba with Adaptive Fusion
di: Meng, Zelin, et al.
Pubblicazione: (2024)
di: Meng, Zelin, et al.
Pubblicazione: (2024)
Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing
di: Kapse, Saarthak, et al.
Pubblicazione: (2025)
di: Kapse, Saarthak, et al.
Pubblicazione: (2025)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
di: Li, Yan, et al.
Pubblicazione: (2024)
di: Li, Yan, et al.
Pubblicazione: (2024)
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
di: Liu, Shanhui, et al.
Pubblicazione: (2025)
di: Liu, Shanhui, et al.
Pubblicazione: (2025)
MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution
di: Chang, Hua, et al.
Pubblicazione: (2025)
di: Chang, Hua, et al.
Pubblicazione: (2025)
Dynamic Vision Mamba
di: Wu, Mengxuan, et al.
Pubblicazione: (2025)
di: Wu, Mengxuan, et al.
Pubblicazione: (2025)
SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging
di: Tran, Nhat Thanh, et al.
Pubblicazione: (2025)
di: Tran, Nhat Thanh, et al.
Pubblicazione: (2025)
Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
Separators in Enhancing Autoregressive Pretraining for Vision Mamba
di: Liu, Hanpeng, et al.
Pubblicazione: (2026)
di: Liu, Hanpeng, et al.
Pubblicazione: (2026)
PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models
di: Meng, Yu, et al.
Pubblicazione: (2025)
di: Meng, Yu, et al.
Pubblicazione: (2025)
ME-Mamba: Multi-Expert Mamba with Efficient Knowledge Capture and Fusion for Multimodal Survival Analysis
di: Zhang, Chengsheng, et al.
Pubblicazione: (2025)
di: Zhang, Chengsheng, et al.
Pubblicazione: (2025)
TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model
di: Yang, Cheng, et al.
Pubblicazione: (2025)
di: Yang, Cheng, et al.
Pubblicazione: (2025)
MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation
di: Rahman, Md Maklachur, et al.
Pubblicazione: (2026)
di: Rahman, Md Maklachur, et al.
Pubblicazione: (2026)
Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection
di: Senadeera, Damith Chamalke, et al.
Pubblicazione: (2025)
di: Senadeera, Damith Chamalke, et al.
Pubblicazione: (2025)
TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
di: Yang, Ziyue, et al.
Pubblicazione: (2026)
di: Yang, Ziyue, et al.
Pubblicazione: (2026)
DeMansia: Mamba Never Forgets Any Tokens
di: Fang, Ricky
Pubblicazione: (2024)
di: Fang, Ricky
Pubblicazione: (2024)
A Survey on Mamba Architecture for Vision Applications
di: Ibrahim, Fady, et al.
Pubblicazione: (2025)
di: Ibrahim, Fady, et al.
Pubblicazione: (2025)
Mamba Fusion: Learning Actions Through Questioning
di: Dong, Zhikang, et al.
Pubblicazione: (2024)
di: Dong, Zhikang, et al.
Pubblicazione: (2024)
Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation
di: Zheng, Anlin, et al.
Pubblicazione: (2025)
di: Zheng, Anlin, et al.
Pubblicazione: (2025)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2026)
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2026)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
OuroMamba: A Data-Free Quantization Framework for Vision Mamba
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
Beyond ZOH: Advanced Discretization Strategies for Vision Mamba
di: Ibrahim, Fady, et al.
Pubblicazione: (2026)
di: Ibrahim, Fady, et al.
Pubblicazione: (2026)
Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
di: Heo, Seongsoo, et al.
Pubblicazione: (2025)
di: Heo, Seongsoo, et al.
Pubblicazione: (2025)
Selective Visual Prompting in Vision Mamba
di: Yao, Yifeng, et al.
Pubblicazione: (2024)
di: Yao, Yifeng, et al.
Pubblicazione: (2024)
SF-Mamba: Rethinking State Space Model for Vision
di: Yoshimura, Masakazu, et al.
Pubblicazione: (2026)
di: Yoshimura, Masakazu, et al.
Pubblicazione: (2026)
FMRFT: Fusion Mamba and DETR for Query Time Sequence Intersection Fish Tracking
di: Yao, Mingyuan, et al.
Pubblicazione: (2024)
di: Yao, Mingyuan, et al.
Pubblicazione: (2024)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
di: Tang, Zicong, et al.
Pubblicazione: (2025)
di: Tang, Zicong, et al.
Pubblicazione: (2025)
Rethinking Token Reduction for Large Vision-Language Models
di: Wang, Yi, et al.
Pubblicazione: (2026)
di: Wang, Yi, et al.
Pubblicazione: (2026)
MambaOut: Do We Really Need Mamba for Vision?
di: Yu, Weihao, et al.
Pubblicazione: (2024)
di: Yu, Weihao, et al.
Pubblicazione: (2024)
CFMD: Dynamic Cross-layer Feature Fusion for Salient Object Detection
di: Lian, Jin, et al.
Pubblicazione: (2025)
di: Lian, Jin, et al.
Pubblicazione: (2025)
TAP-SLF: Parameter-Efficient Adaptation of Vision Foundation Models for Multi-Task Ultrasound Image Analysis
di: Wan, Hui, et al.
Pubblicazione: (2026)
di: Wan, Hui, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation
di: Liu, Che, et al.
Pubblicazione: (2024) -
Fusion-Mamba for Cross-modality Object Detection
di: Dong, Wenhao, et al.
Pubblicazione: (2024) -
EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion
di: Yang, Zichuan, et al.
Pubblicazione: (2025) -
COMO: Cross-Mamba Interaction and Offset-Guided Fusion for Multimodal Object Detection
di: Liu, Chang, et al.
Pubblicazione: (2024) -
SwiftVLM: Efficient Vision-Language Model Inference via Cross-Layer Token Bypass
di: Qian, Chen, et al.
Pubblicazione: (2026)