InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zongyi, Zhao, Junchuan, Lee, Francis Bu Sung, Yee, Andrew Zi Han |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
von: Zeng, Wei, et al.
Veröffentlicht: (2025)
von: Zeng, Wei, et al.
Veröffentlicht: (2025)
Memo2496: Expert-Annotated Dataset and Dual-View Adaptive Framework for Music Emotion Recognition
von: Li, Qilin, et al.
Veröffentlicht: (2025)
von: Li, Qilin, et al.
Veröffentlicht: (2025)
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
von: Xu, Xinmeng, et al.
Veröffentlicht: (2026)
von: Xu, Xinmeng, et al.
Veröffentlicht: (2026)
TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation
von: Lee, Yubeen, et al.
Veröffentlicht: (2025)
von: Lee, Yubeen, et al.
Veröffentlicht: (2025)
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
Stage-Adaptive Reliability Modeling for Continuous Valence-Arousal Estimation
von: Lee, Yubeen, et al.
Veröffentlicht: (2026)
von: Lee, Yubeen, et al.
Veröffentlicht: (2026)
Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition
von: Koo, Inyong, et al.
Veröffentlicht: (2026)
von: Koo, Inyong, et al.
Veröffentlicht: (2026)
A Survey on Multimodal Music Emotion Recognition
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
CLAIP-Emo: Parameter-Efficient Adaptation of Language-supervised models for In-the-Wild Audiovisual Emotion Recognition
von: Chen, Yin, et al.
Veröffentlicht: (2025)
von: Chen, Yin, et al.
Veröffentlicht: (2025)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
von: Zhou, Dingkun, et al.
Veröffentlicht: (2025)
von: Zhou, Dingkun, et al.
Veröffentlicht: (2025)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
AMB-DSGDN: Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network for Multimodal Emotion Recognition
von: Wang, Yunsheng, et al.
Veröffentlicht: (2026)
von: Wang, Yunsheng, et al.
Veröffentlicht: (2026)
MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
von: Yang, Meng, et al.
Veröffentlicht: (2026)
von: Yang, Meng, et al.
Veröffentlicht: (2026)
MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
von: Shao, Yongqi, et al.
Veröffentlicht: (2025)
von: Shao, Yongqi, et al.
Veröffentlicht: (2025)
Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
von: Shen, Nanhan, et al.
Veröffentlicht: (2026)
von: Shen, Nanhan, et al.
Veröffentlicht: (2026)
SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
von: Desai, Shail, et al.
Veröffentlicht: (2025)
von: Desai, Shail, et al.
Veröffentlicht: (2025)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation
von: Qiu, Ke, et al.
Veröffentlicht: (2026)
von: Qiu, Ke, et al.
Veröffentlicht: (2026)
A Unified Framework for Modality-Agnostic Deepfakes Detection
von: Yu, Cai, et al.
Veröffentlicht: (2023)
von: Yu, Cai, et al.
Veröffentlicht: (2023)
SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning
von: Li, Zhu, et al.
Veröffentlicht: (2026)
von: Li, Zhu, et al.
Veröffentlicht: (2026)
Personality-Enhanced Multimodal Depression Detection in the Elderly
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
Gesture2Music: A Low-Latency Real-Time Framework for Continuous Gesture-Driven Music Generation
von: Jeyaraj, Rathinaraja, et al.
Veröffentlicht: (2025)
von: Jeyaraj, Rathinaraja, et al.
Veröffentlicht: (2025)
MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
von: Pan, Zihan, et al.
Veröffentlicht: (2025)
Ensembling Synchronisation-based and Face-Voice Association Paradigms for Robust Active Speaker Detection in Egocentric Recordings
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
Audio-Visual Separation with Hierarchical Fusion and Representation Alignment
von: Hu, Han, et al.
Veröffentlicht: (2025)
von: Hu, Han, et al.
Veröffentlicht: (2025)
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
von: Li, Fu, et al.
Veröffentlicht: (2025)
von: Li, Fu, et al.
Veröffentlicht: (2025)
MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
EMO100DB: An Open Dataset of Improvised Songs with Emotion Data
von: Hwang, Daeun, et al.
Veröffentlicht: (2025)
von: Hwang, Daeun, et al.
Veröffentlicht: (2025)
Can We Hear from Events? Generating Speech from Event Camera
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
MF-AED-AEC: Speech Emotion Recognition by Leveraging Multimodal Fusion, Asr Error Detection, and Asr Error Correction
von: He, Jiajun, et al.
Veröffentlicht: (2024)
von: He, Jiajun, et al.
Veröffentlicht: (2024)
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
von: Ma, Ziyang, et al.
Veröffentlicht: (2026)
von: Ma, Ziyang, et al.
Veröffentlicht: (2026)
Research on Piano Timbre Transformation System Based on Diffusion Model
von: Hsu, Chun-Chieh, et al.
Veröffentlicht: (2026)
von: Hsu, Chun-Chieh, et al.
Veröffentlicht: (2026)
Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
von: Xia, Haiying, et al.
Veröffentlicht: (2025)
von: Xia, Haiying, et al.
Veröffentlicht: (2025)
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
Dance-to-Music Generation with Encoder-based Textual Inversion
von: Li, Sifei, et al.
Veröffentlicht: (2024)
von: Li, Sifei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
von: Zeng, Wei, et al.
Veröffentlicht: (2025) -
Memo2496: Expert-Annotated Dataset and Dual-View Adaptive Framework for Music Emotion Recognition
von: Li, Qilin, et al.
Veröffentlicht: (2025) -
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
von: Xu, Xinmeng, et al.
Veröffentlicht: (2026) -
TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation
von: Lee, Yubeen, et al.
Veröffentlicht: (2025) -
CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face Generation
von: Wu, Kangyi, et al.
Veröffentlicht: (2025)