Mamba-Enhanced Text-Audio-Video Alignment Network for Emotion Recognition in Conversations
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinran, Fan, Xiaomao, Wu, Qingyang, Peng, Xiaojiang, Li, Ye |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context Transformer
by: Li, Xinpeng, et al.
Published: (2024)
by: Li, Xinpeng, et al.
Published: (2024)
LEAF: Unveiling Two Sides of the Same Coin in Semi-supervised Facial Expression Recognition
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
by: Kim, Yeonju, et al.
Published: (2024)
by: Kim, Yeonju, et al.
Published: (2024)
MIMIC: Mask Image Pre-training with Mix Contrastive Fine-tuning for Facial Expression Recognition
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2025)
by: Wang, Langyu, et al.
Published: (2025)
DualMamba: A Lightweight Spectral-Spatial Mamba-Convolution Network for Hyperspectral Image Classification
by: Sheng, Jiamu, et al.
Published: (2024)
by: Sheng, Jiamu, et al.
Published: (2024)
Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
by: Shi, Tong, et al.
Published: (2024)
by: Shi, Tong, et al.
Published: (2024)
MambaPlace:Text-to-Point-Cloud Cross-Modal Place Recognition with Attention Mamba Mechanisms
by: Shang, Tianyi, et al.
Published: (2024)
by: Shang, Tianyi, et al.
Published: (2024)
Stain-aware Domain Alignment for Imbalance Blood Cell Classification
by: Li, Yongcheng, et al.
Published: (2024)
by: Li, Yongcheng, et al.
Published: (2024)
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
Low-Contrast-Enhanced Contrastive Learning for Semi-Supervised Endoscopic Image Segmentation
by: Cai, Lingcong, et al.
Published: (2024)
by: Cai, Lingcong, et al.
Published: (2024)
Hear the Scene: Audio-Enhanced Text Spotting
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
3D Landmark Detection on Human Point Clouds: A Benchmark and A Dual Cascade Point Transformer Framework
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
Enhancing Micro Gesture Recognition for Emotion Understanding via Context-aware Visual-Text Contrastive Learning
by: Li, Deng, et al.
Published: (2024)
by: Li, Deng, et al.
Published: (2024)
RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
SAM-FNet: SAM-Guided Fusion Network for Laryngo-Pharyngeal Tumor Detection
by: Wei, Jia, et al.
Published: (2024)
by: Wei, Jia, et al.
Published: (2024)
Dual-path Collaborative Generation Network for Emotional Video Captioning
by: Ye, Cheng, et al.
Published: (2024)
by: Ye, Cheng, et al.
Published: (2024)
VisionLLM-based Multimodal Fusion Network for Glottic Carcinoma Early Detection
by: Jin, Zhaohui, et al.
Published: (2024)
by: Jin, Zhaohui, et al.
Published: (2024)
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
by: Li, Qifei, et al.
Published: (2024)
by: Li, Qifei, et al.
Published: (2024)
Dynamic Vision Mamba
by: Wu, Mengxuan, et al.
Published: (2025)
by: Wu, Mengxuan, et al.
Published: (2025)
3D-LSPTM: An Automatic Framework with 3D-Large-Scale Pretrained Model for Laryngeal Cancer Detection Using Laryngoscopic Videos
by: Qiu, Meiyu, et al.
Published: (2024)
by: Qiu, Meiyu, et al.
Published: (2024)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D Retrieval
by: Li, Wenrui, et al.
Published: (2025)
by: Li, Wenrui, et al.
Published: (2025)
Storyboard guided Alignment for Fine-grained Video Action Recognition
by: Liu, Enqi, et al.
Published: (2024)
by: Liu, Enqi, et al.
Published: (2024)
Video Emotion Open-vocabulary Recognition Based on Multimodal Large Language Model
by: Ge, Mengying, et al.
Published: (2024)
by: Ge, Mengying, et al.
Published: (2024)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2025)
by: Liu, Yunze, et al.
Published: (2025)
Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-Sequence
by: Huang, Wenbo, et al.
Published: (2024)
by: Huang, Wenbo, et al.
Published: (2024)
Multi-modal Speech Emotion Recognition via Feature Distribution Adaptation Network
by: Li, Shaokai, et al.
Published: (2024)
by: Li, Shaokai, et al.
Published: (2024)
Hierarchical Disentanglement-Alignment Network for Robust SAR Vehicle Recognition
by: Li, Weijie, et al.
Published: (2023)
by: Li, Weijie, et al.
Published: (2023)
Knowledge-Enhanced Facial Expression Recognition with Emotional-to-Neutral Transformation
by: Li, Hangyu, et al.
Published: (2024)
by: Li, Hangyu, et al.
Published: (2024)
ICANet: A Method of Short Video Emotion Recognition Driven by Multimodal Data
by: Wu, Xuecheng, et al.
Published: (2022)
by: Wu, Xuecheng, et al.
Published: (2022)
VideoMamba: State Space Model for Efficient Video Understanding
by: Li, Kunchang, et al.
Published: (2024)
by: Li, Kunchang, et al.
Published: (2024)
Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba
by: Wang, Yuhang, et al.
Published: (2025)
by: Wang, Yuhang, et al.
Published: (2025)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
by: Huang, Dawei, et al.
Published: (2025)
by: Huang, Dawei, et al.
Published: (2025)
MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition
by: Li, Feng, et al.
Published: (2025)
by: Li, Feng, et al.
Published: (2025)
eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos
by: Wu, Xuecheng, et al.
Published: (2025)
by: Wu, Xuecheng, et al.
Published: (2025)
Similar Items
-
Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context Transformer
by: Li, Xinpeng, et al.
Published: (2024) -
LEAF: Unveiling Two Sides of the Same Coin in Semi-supervised Facial Expression Recognition
by: Zhang, Fan, et al.
Published: (2024) -
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
by: Kim, Yeonju, et al.
Published: (2024) -
MIMIC: Mask Image Pre-training with Mix Contrastive Fine-tuning for Facial Expression Recognition
by: Zhang, Fan, et al.
Published: (2024) -
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2025)