Early Joint Learning of Emotion Information Makes MultiModal Model Understand You Better
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ge, Mengying, Li, Mingyang, Tang, Dongkai, Li, Pengbo, Liu, Kuo, Deng, Shuhao, Pu, Songbai, Liu, Long, Song, Yang, Zhang, Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
A Unified Framework for Modality-Agnostic Deepfakes Detection
von: Yu, Cai, et al.
Veröffentlicht: (2023)
von: Yu, Cai, et al.
Veröffentlicht: (2023)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2024)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
A Survey on Cross-Modal Interaction Between Music and Multimodal Data
von: Li, Sifei, et al.
Veröffentlicht: (2025)
von: Li, Sifei, et al.
Veröffentlicht: (2025)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Music Emotion Recognition
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
Emotion-Aligned Contrastive Learning Between Images and Music
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
von: Lei, Ke, et al.
Veröffentlicht: (2026)
von: Lei, Ke, et al.
Veröffentlicht: (2026)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
A Study on Synthesizing Expressive Violin Performances: Approaches and Comparisons
von: Hung, Tzu-Yun, et al.
Veröffentlicht: (2024)
von: Hung, Tzu-Yun, et al.
Veröffentlicht: (2024)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
von: Liu, Shengqiang, et al.
Veröffentlicht: (2024)
von: Liu, Shengqiang, et al.
Veröffentlicht: (2024)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
von: Wang, Anna, et al.
Veröffentlicht: (2024)
von: Wang, Anna, et al.
Veröffentlicht: (2024)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
Separate Anything You Describe
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
von: Cui, Meng, et al.
Veröffentlicht: (2023)
von: Cui, Meng, et al.
Veröffentlicht: (2023)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
Personality-Enhanced Multimodal Depression Detection in the Elderly
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
Efficient Video to Audio Mapper with Visual Scene Detection
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
A Survey of Foundation Models for Music Understanding
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
Visual-based spatial audio generation system for multi-speaker environments
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
Intelligent Text-Conditioned Music Generation
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
von: Rong, Yan, et al.
Veröffentlicht: (2025) -
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
von: Deng, Jiajun, et al.
Veröffentlicht: (2025) -
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026) -
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025) -
A Unified Framework for Modality-Agnostic Deepfakes Detection
von: Yu, Cai, et al.
Veröffentlicht: (2023)