Mamba Fusion: Learning Actions Through Questioning
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Zhikang, Beedu, Apoorva, Sheinkopf, Jason, Essa, Irfan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HierSum: A Global and Local Attention Mechanism for Video Summarization
by: Beedu, Apoorva, et al.
Published: (2025)
by: Beedu, Apoorva, et al.
Published: (2025)
Exploring Efficient Foundational Multi-modal Models for Video Summarization
by: Samel, Karan, et al.
Published: (2024)
by: Samel, Karan, et al.
Published: (2024)
On the Efficacy of Text-Based Input Modalities for Action Anticipation
by: Beedu, Apoorva, et al.
Published: (2024)
by: Beedu, Apoorva, et al.
Published: (2024)
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
by: Haresamudram, Harish, et al.
Published: (2024)
by: Haresamudram, Harish, et al.
Published: (2024)
Fusion-Mamba for Cross-modality Object Detection
by: Dong, Wenhao, et al.
Published: (2024)
by: Dong, Wenhao, et al.
Published: (2024)
DepthMamba with Adaptive Fusion
by: Meng, Zelin, et al.
Published: (2024)
by: Meng, Zelin, et al.
Published: (2024)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
by: Marmon, Andrew, et al.
Published: (2024)
by: Marmon, Andrew, et al.
Published: (2024)
COMO: Cross-Mamba Interaction and Offset-Guided Fusion for Multimodal Object Detection
by: Liu, Chang, et al.
Published: (2024)
by: Liu, Chang, et al.
Published: (2024)
EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion
by: Yang, Zichuan, et al.
Published: (2025)
by: Yang, Zichuan, et al.
Published: (2025)
SpikMamba: When SNN meets Mamba in Event-based Human Action Recognition
by: Chen, Jiaqi, et al.
Published: (2024)
by: Chen, Jiaqi, et al.
Published: (2024)
ME-Mamba: Multi-Expert Mamba with Efficient Knowledge Capture and Fusion for Multimodal Survival Analysis
by: Zhang, Chengsheng, et al.
Published: (2025)
by: Zhang, Chengsheng, et al.
Published: (2025)
SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
by: Zhao, Ruosen, et al.
Published: (2025)
by: Zhao, Ruosen, et al.
Published: (2025)
WDFFU-Mamba: A Wavelet-guided Dual-attention Feature Fusion Mamba for Breast Tumor Segmentation in Ultrasound Images
by: Cai, Guoping, et al.
Published: (2025)
by: Cai, Guoping, et al.
Published: (2025)
Famba-V: Fast Vision Mamba with Cross-Layer Token Fusion
by: Shen, Hui, et al.
Published: (2024)
by: Shen, Hui, et al.
Published: (2024)
FMRFT: Fusion Mamba and DETR for Query Time Sequence Intersection Fish Tracking
by: Yao, Mingyuan, et al.
Published: (2024)
by: Yao, Mingyuan, et al.
Published: (2024)
MDDFNet: Mamba-based Dynamic Dual Fusion Network for Traffic Sign Detection
by: Yu, TianYi
Published: (2025)
by: Yu, TianYi
Published: (2025)
CluMo: Cluster-based Modality Fusion Prompt for Continual Learning in Visual Question Answering
by: Cai, Yuliang, et al.
Published: (2024)
by: Cai, Yuliang, et al.
Published: (2024)
MDF-MLLM: Deep Fusion Through Cross-Modal Feature Alignment for Contextually Aware Fundoscopic Image Classification
by: Jordan, Jason, et al.
Published: (2025)
by: Jordan, Jason, et al.
Published: (2025)
MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation
by: Rahman, Md Maklachur, et al.
Published: (2026)
by: Rahman, Md Maklachur, et al.
Published: (2026)
LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
DEPFusion: Dual-Domain Enhancement and Priority-Guided Mamba Fusion for UAV Multispectral Object Detection
by: Li, Shucong, et al.
Published: (2025)
by: Li, Shucong, et al.
Published: (2025)
MambaOVSR: Multiscale Fusion with Global Motion Modeling for Chinese Opera Video Super-Resolution
by: Chang, Hua, et al.
Published: (2025)
by: Chang, Hua, et al.
Published: (2025)
MambaTAD: When State-Space Models Meet Long-Range Temporal Action Detection
by: Lu, Hui, et al.
Published: (2025)
by: Lu, Hui, et al.
Published: (2025)
Every Image Listens, Every Image Dances: Music-Driven Image Animation
by: Dong, Zhikang, et al.
Published: (2025)
by: Dong, Zhikang, et al.
Published: (2025)
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
by: Zhang, Junkai, et al.
Published: (2025)
by: Zhang, Junkai, et al.
Published: (2025)
TextMamba: Scene Text Detector with Mamba
by: Zhao, Qiyan, et al.
Published: (2025)
by: Zhao, Qiyan, et al.
Published: (2025)
BioFusionNet: Deep Learning-Based Survival Risk Stratification in ER+ Breast Cancer Through Multifeature and Multimodal Data Fusion
by: Mondol, Raktim Kumar, et al.
Published: (2024)
by: Mondol, Raktim Kumar, et al.
Published: (2024)
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
by: Liu, Shanhui, et al.
Published: (2025)
by: Liu, Shanhui, et al.
Published: (2025)
CellMamba: Adaptive Mamba for Accurate and Efficient Cell Detection
by: Liu, Ruochen, et al.
Published: (2025)
by: Liu, Ruochen, et al.
Published: (2025)
Feature Fusion for Human Activity Recognition using Parameter-Optimized Multi-Stage Graph Convolutional Network and Transformer Models
by: Belal, Mohammad, et al.
Published: (2024)
by: Belal, Mohammad, et al.
Published: (2024)
Q-FSRU: Quantum-Augmented Frequency-Spectral Fusion for Medical Visual Question Answering
by: Thakur, Rakesh, et al.
Published: (2025)
by: Thakur, Rakesh, et al.
Published: (2025)
OuroMamba: A Data-Free Quantization Framework for Vision Mamba
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images
by: Zhang, Yunfei, et al.
Published: (2025)
by: Zhang, Yunfei, et al.
Published: (2025)
DMS2F-HAD: A Dual-branch Mamba-based Spatial-Spectral Fusion Network for Hyperspectral Anomaly Detection
by: Pant, Aayushma, et al.
Published: (2026)
by: Pant, Aayushma, et al.
Published: (2026)
FMaMIL: Frequency-Driven Mamba Multi-Instance Learning for Weakly Supervised Lesion Segmentation in Medical Images
by: Cheng, Hangbei, et al.
Published: (2025)
by: Cheng, Hangbei, et al.
Published: (2025)
Dynamic Vision Mamba
by: Wu, Mengxuan, et al.
Published: (2025)
by: Wu, Mengxuan, et al.
Published: (2025)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
by: Anaissi, Ali, et al.
Published: (2025)
by: Anaissi, Ali, et al.
Published: (2025)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
by: Kwon, Minchan, et al.
Published: (2026)
by: Kwon, Minchan, et al.
Published: (2026)
CORE-ReID V2: Advancing the Domain Adaptation for Object Re-Identification with Optimized Training and Ensemble Fusion
by: Nguyen, Trinh Quoc, et al.
Published: (2025)
by: Nguyen, Trinh Quoc, et al.
Published: (2025)
Similar Items
-
HierSum: A Global and Local Attention Mechanism for Video Summarization
by: Beedu, Apoorva, et al.
Published: (2025) -
Exploring Efficient Foundational Multi-modal Models for Video Summarization
by: Samel, Karan, et al.
Published: (2024) -
On the Efficacy of Text-Based Input Modalities for Action Anticipation
by: Beedu, Apoorva, et al.
Published: (2024) -
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
by: Haresamudram, Harish, et al.
Published: (2024) -
Fusion-Mamba for Cross-modality Object Detection
by: Dong, Wenhao, et al.
Published: (2024)