AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gong, Sitong, Zhuge, Yunzhi, Zhang, Lu, Wang, Yifan, Zhang, Pingping, Wang, Lijun, Lu, Huchuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Complementary and Contrastive Learning for Audio-Visual Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
Parameter Aware Mamba Model for Multi-task Dense Prediction
von: Yu, Xinzhuo, et al.
Veröffentlicht: (2025)
von: Yu, Xinzhuo, et al.
Veröffentlicht: (2025)
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
von: Zhuge, Yunzhi, et al.
Veröffentlicht: (2025)
von: Zhuge, Yunzhi, et al.
Veröffentlicht: (2025)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
Learning Universal Features for Generalizable Image Forgery Localization
von: Zhao, Hengrun, et al.
Veröffentlicht: (2025)
von: Zhao, Hengrun, et al.
Veröffentlicht: (2025)
Spatial-Frequency Enhanced Mamba for Multi-Modal Image Fusion
von: Sun, Hui, et al.
Veröffentlicht: (2025)
von: Sun, Hui, et al.
Veröffentlicht: (2025)
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
von: Ye, Chengyang, et al.
Veröffentlicht: (2024)
von: Ye, Chengyang, et al.
Veröffentlicht: (2024)
NeuroMamba: Multi-Perspective Feature Interaction with Visual Mamba for Neuron Segmentation
von: Jiang, Liuyun, et al.
Veröffentlicht: (2026)
von: Jiang, Liuyun, et al.
Veröffentlicht: (2026)
IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
von: Zhu, Yixin, et al.
Veröffentlicht: (2026)
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
von: Wang, Yaoting, et al.
Veröffentlicht: (2024)
von: Wang, Yaoting, et al.
Veröffentlicht: (2024)
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
von: Zhang, Lu, et al.
Veröffentlicht: (2025)
von: Zhang, Lu, et al.
Veröffentlicht: (2025)
DefMamba: Deformable Visual State Space Model
von: Liu, Leiye, et al.
Veröffentlicht: (2025)
von: Liu, Leiye, et al.
Veröffentlicht: (2025)
Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling
von: Feng, Xiyan, et al.
Veröffentlicht: (2026)
von: Feng, Xiyan, et al.
Veröffentlicht: (2026)
A Survey on Visual Mamba
von: Zhang, Hanwei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanwei, et al.
Veröffentlicht: (2024)
Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation
von: Wan, Zifu, et al.
Veröffentlicht: (2024)
von: Wan, Zifu, et al.
Veröffentlicht: (2024)
P-Mamba: Marrying Perona Malik Diffusion with Mamba for Efficient Pediatric Echocardiographic Left Ventricular Segmentation
von: Ye, Zi, et al.
Veröffentlicht: (2024)
von: Ye, Zi, et al.
Veröffentlicht: (2024)
MambaVT: Spatio-Temporal Contextual Modeling for robust RGB-T Tracking
von: Lai, Simiao, et al.
Veröffentlicht: (2024)
von: Lai, Simiao, et al.
Veröffentlicht: (2024)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
von: He, Junwen, et al.
Veröffentlicht: (2024)
von: He, Junwen, et al.
Veröffentlicht: (2024)
Other Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification
von: Wang, Yingquan, et al.
Veröffentlicht: (2024)
von: Wang, Yingquan, et al.
Veröffentlicht: (2024)
Multi-Scale and Detail-Enhanced Segment Anything Model for Salient Object Detection
von: Gao, Shixuan, et al.
Veröffentlicht: (2024)
von: Gao, Shixuan, et al.
Veröffentlicht: (2024)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
von: Zhang, Wenbo, et al.
Veröffentlicht: (2024)
von: Zhang, Wenbo, et al.
Veröffentlicht: (2024)
MobileMamba: Lightweight Multi-Receptive Visual Mamba Network
von: He, Haoyang, et al.
Veröffentlicht: (2024)
von: He, Haoyang, et al.
Veröffentlicht: (2024)
MambaDFuse: A Mamba-based Dual-phase Model for Multi-modality Image Fusion
von: Li, Zhe, et al.
Veröffentlicht: (2024)
von: Li, Zhe, et al.
Veröffentlicht: (2024)
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
von: Xiong, Haomiao, et al.
Veröffentlicht: (2025)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
BEV-IO: Enhancing Bird's-Eye-View 3D Detection with Instance Occupancy
von: Zhang, Zaibin, et al.
Veröffentlicht: (2023)
von: Zhang, Zaibin, et al.
Veröffentlicht: (2023)
3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum
von: Zhang, Yuliang, et al.
Veröffentlicht: (2026)
von: Zhang, Yuliang, et al.
Veröffentlicht: (2026)
Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAM
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
AVS-Net: Audio-Visual Scale Net for Self-supervised Monocular Metric Depth Estimation
von: Liu, Xiaohu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaohu, et al.
Veröffentlicht: (2024)
LATex: Leveraging Attribute-based Text Knowledge for Aerial-Ground Person Re-Identification
von: Zhang, Pingping, et al.
Veröffentlicht: (2025)
von: Zhang, Pingping, et al.
Veröffentlicht: (2025)
X-ReID: Multi-granularity Information Interaction for Video-Based Visible-Infrared Person Re-Identification
von: Yu, Chenyang, et al.
Veröffentlicht: (2025)
von: Yu, Chenyang, et al.
Veröffentlicht: (2025)
GlobalMamba: Global Image Serialization for Vision Mamba
von: Wang, Chengkun, et al.
Veröffentlicht: (2024)
von: Wang, Chengkun, et al.
Veröffentlicht: (2024)
StableIdentity: Inserting Anybody into Anywhere at First Sight
von: Wang, Qinghe, et al.
Veröffentlicht: (2024)
von: Wang, Qinghe, et al.
Veröffentlicht: (2024)
Mamba-UNet: UNet-Like Pure Visual Mamba for Medical Image Segmentation
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
von: Wang, Ziyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Complementary and Contrastive Learning for Audio-Visual Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025) -
Parameter Aware Mamba Model for Multi-task Dense Prediction
von: Yu, Xinzhuo, et al.
Veröffentlicht: (2025) -
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025) -
Reinforcing Video Reasoning Segmentation to Think Before It Segments
von: Gong, Sitong, et al.
Veröffentlicht: (2025) -
Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
von: Zhuge, Yunzhi, et al.
Veröffentlicht: (2025)