Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zhixian, Tian, Wenjie, Xie, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts
by: Zhong, Guowei, et al.
Published: (2025)
by: Zhong, Guowei, et al.
Published: (2025)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
by: Zhao, Pengcheng, et al.
Published: (2024)
by: Zhao, Pengcheng, et al.
Published: (2024)
Calibrating Multimodal Consensus for Emotion Recognition
by: Zhong, Guowei, et al.
Published: (2025)
by: Zhong, Guowei, et al.
Published: (2025)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
by: An, Wenbin, et al.
Published: (2025)
by: An, Wenbin, et al.
Published: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
by: Jiang, Chaoya, et al.
Published: (2024)
by: Jiang, Chaoya, et al.
Published: (2024)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
by: Tan, Weiting, et al.
Published: (2025)
by: Tan, Weiting, et al.
Published: (2025)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
Audio-visual training for improved grounding in video-text LLMs
by: Sagare, Shivprasad, et al.
Published: (2024)
by: Sagare, Shivprasad, et al.
Published: (2024)
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
by: Chao, Jianghan, et al.
Published: (2025)
by: Chao, Jianghan, et al.
Published: (2025)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
by: Zhu, Lingsi, et al.
Published: (2026)
by: Zhu, Lingsi, et al.
Published: (2026)
Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
by: Lin, Ronghao, et al.
Published: (2025)
by: Lin, Ronghao, et al.
Published: (2025)
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
by: Astrid, Marcella, et al.
Published: (2024)
by: Astrid, Marcella, et al.
Published: (2024)
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Noise-Tolerant Learning for Audio-Visual Action Recognition
by: Han, Haochen, et al.
Published: (2022)
by: Han, Haochen, et al.
Published: (2022)
Beyond Coarse-Grained Matching in Video-Text Retrieval
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
by: Jiang, Xin, et al.
Published: (2024)
by: Jiang, Xin, et al.
Published: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025)
by: Yan, Zehong, et al.
Published: (2025)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
by: Jiang, Jingjing, et al.
Published: (2025)
by: Jiang, Jingjing, et al.
Published: (2025)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
by: Wang, Bing, et al.
Published: (2024)
by: Wang, Bing, et al.
Published: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
by: Bin, Yi, et al.
Published: (2024)
by: Bin, Yi, et al.
Published: (2024)
Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited Samples
by: Fang, Ziye, et al.
Published: (2023)
by: Fang, Ziye, et al.
Published: (2023)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
by: Jiang, Ruixiang, et al.
Published: (2025)
by: Jiang, Ruixiang, et al.
Published: (2025)
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
by: Satar, Burak, et al.
Published: (2025)
by: Satar, Burak, et al.
Published: (2025)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
by: Wang, Kangsheng, et al.
Published: (2025)
by: Wang, Kangsheng, et al.
Published: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
by: Gan, Chengguang, et al.
Published: (2025)
by: Gan, Chengguang, et al.
Published: (2025)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
by: Masumura, Ryo, et al.
Published: (2025)
by: Masumura, Ryo, et al.
Published: (2025)
HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
by: Yang, Yiqing, et al.
Published: (2025)
by: Yang, Yiqing, et al.
Published: (2025)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
by: Wang, Bing, et al.
Published: (2025)
by: Wang, Bing, et al.
Published: (2025)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
by: Guo, Xinyue, et al.
Published: (2025)
by: Guo, Xinyue, et al.
Published: (2025)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
by: Araujo, Edson, et al.
Published: (2026)
by: Araujo, Edson, et al.
Published: (2026)
Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
by: Caffagni, Davide, et al.
Published: (2025)
by: Caffagni, Davide, et al.
Published: (2025)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
by: Ghaleb, Esam, et al.
Published: (2025)
by: Ghaleb, Esam, et al.
Published: (2025)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
by: Wu, Jiaying, et al.
Published: (2025)
by: Wu, Jiaying, et al.
Published: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
by: Luo, Chuwei, et al.
Published: (2022)
by: Luo, Chuwei, et al.
Published: (2022)
Similar Items
-
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
by: Cheng, Zebang, et al.
Published: (2024) -
Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts
by: Zhong, Guowei, et al.
Published: (2025) -
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
by: Zhao, Pengcheng, et al.
Published: (2024) -
Calibrating Multimodal Consensus for Emotion Recognition
by: Zhong, Guowei, et al.
Published: (2025) -
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
by: An, Wenbin, et al.
Published: (2025)