Salvato in:
| Autori principali: | Ge, Mengying, Li, Mingyang, Tang, Dongkai, Li, Pengbo, Liu, Kuo, Deng, Shuhao, Pu, Songbai, Liu, Long, Song, Yang, Zhang, Tao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2409.18971 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
di: Rong, Yan, et al.
Pubblicazione: (2025)
di: Rong, Yan, et al.
Pubblicazione: (2025)
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
di: Deng, Jiajun, et al.
Pubblicazione: (2025)
di: Deng, Jiajun, et al.
Pubblicazione: (2025)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
di: Su, Fei, et al.
Pubblicazione: (2026)
di: Su, Fei, et al.
Pubblicazione: (2026)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
di: Lin, Yueqian, et al.
Pubblicazione: (2025)
di: Lin, Yueqian, et al.
Pubblicazione: (2025)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
di: Deng, Zeyu, et al.
Pubblicazione: (2025)
di: Deng, Zeyu, et al.
Pubblicazione: (2025)
A Unified Framework for Modality-Agnostic Deepfakes Detection
di: Yu, Cai, et al.
Pubblicazione: (2023)
di: Yu, Cai, et al.
Pubblicazione: (2023)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
di: Cong, Gaoxiang, et al.
Pubblicazione: (2024)
di: Cong, Gaoxiang, et al.
Pubblicazione: (2024)
A Survey on Cross-Modal Interaction Between Music and Multimodal Data
di: Li, Sifei, et al.
Pubblicazione: (2025)
di: Li, Sifei, et al.
Pubblicazione: (2025)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
di: Wang, Haoxu, et al.
Pubblicazione: (2024)
di: Wang, Haoxu, et al.
Pubblicazione: (2024)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
di: Tian, Wenjie, et al.
Pubblicazione: (2025)
di: Tian, Wenjie, et al.
Pubblicazione: (2025)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
di: Liu, Haohe, et al.
Pubblicazione: (2024)
di: Liu, Haohe, et al.
Pubblicazione: (2024)
A Survey on Multimodal Music Emotion Recognition
di: Liyanarachchi, Rashini, et al.
Pubblicazione: (2025)
di: Liyanarachchi, Rashini, et al.
Pubblicazione: (2025)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
di: Lei, Ke, et al.
Pubblicazione: (2026)
di: Lei, Ke, et al.
Pubblicazione: (2026)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
di: Liu, Shansong, et al.
Pubblicazione: (2023)
di: Liu, Shansong, et al.
Pubblicazione: (2023)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
di: Liu, Shansong, et al.
Pubblicazione: (2024)
di: Liu, Shansong, et al.
Pubblicazione: (2024)
Emotion-Aligned Contrastive Learning Between Images and Music
di: Stewart, Shanti, et al.
Pubblicazione: (2023)
di: Stewart, Shanti, et al.
Pubblicazione: (2023)
A Study on Synthesizing Expressive Violin Performances: Approaches and Comparisons
di: Hung, Tzu-Yun, et al.
Pubblicazione: (2024)
di: Hung, Tzu-Yun, et al.
Pubblicazione: (2024)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
di: Li, Xiaolou, et al.
Pubblicazione: (2024)
di: Li, Xiaolou, et al.
Pubblicazione: (2024)
Separate Anything You Describe
di: Liu, Xubo, et al.
Pubblicazione: (2023)
di: Liu, Xubo, et al.
Pubblicazione: (2023)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
di: Zhang, Xiaohui, et al.
Pubblicazione: (2024)
di: Zhang, Xiaohui, et al.
Pubblicazione: (2024)
DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
di: Liu, Shengqiang, et al.
Pubblicazione: (2024)
di: Liu, Shengqiang, et al.
Pubblicazione: (2024)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
di: Wang, Anna, et al.
Pubblicazione: (2024)
di: Wang, Anna, et al.
Pubblicazione: (2024)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
di: He, Jianfeng, et al.
Pubblicazione: (2023)
di: He, Jianfeng, et al.
Pubblicazione: (2023)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
di: Sudarsanam, Parthasaarathy, et al.
Pubblicazione: (2025)
di: Sudarsanam, Parthasaarathy, et al.
Pubblicazione: (2025)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
di: Guo, Hongming, et al.
Pubblicazione: (2024)
di: Guo, Hongming, et al.
Pubblicazione: (2024)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
di: Guo, Zixin, et al.
Pubblicazione: (2024)
di: Guo, Zixin, et al.
Pubblicazione: (2024)
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
di: Li, Yuanchao, et al.
Pubblicazione: (2024)
Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
di: Kim, Minyoung, et al.
Pubblicazione: (2025)
di: Kim, Minyoung, et al.
Pubblicazione: (2025)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
di: Cui, Meng, et al.
Pubblicazione: (2023)
di: Cui, Meng, et al.
Pubblicazione: (2023)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
di: Liu, Qianhui, et al.
Pubblicazione: (2024)
di: Liu, Qianhui, et al.
Pubblicazione: (2024)
A Survey of Foundation Models for Music Understanding
di: Li, Wenjun, et al.
Pubblicazione: (2024)
di: Li, Wenjun, et al.
Pubblicazione: (2024)
Personality-Enhanced Multimodal Depression Detection in the Elderly
di: Wang, Honghong, et al.
Pubblicazione: (2025)
di: Wang, Honghong, et al.
Pubblicazione: (2025)
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
di: Zhao, Qihao, et al.
Pubblicazione: (2026)
di: Zhao, Qihao, et al.
Pubblicazione: (2026)
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
di: Liu, Renhang, et al.
Pubblicazione: (2024)
di: Liu, Renhang, et al.
Pubblicazione: (2024)
M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
di: Li, Yupei, et al.
Pubblicazione: (2024)
di: Li, Yupei, et al.
Pubblicazione: (2024)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
di: Li, Maomao, et al.
Pubblicazione: (2026)
di: Li, Maomao, et al.
Pubblicazione: (2026)
Efficient Video to Audio Mapper with Visual Scene Detection
di: Yi, Mingjing, et al.
Pubblicazione: (2024)
di: Yi, Mingjing, et al.
Pubblicazione: (2024)
An automatic mixing speech enhancement system for multi-track audio
di: Liu, Xiaojing, et al.
Pubblicazione: (2024)
di: Liu, Xiaojing, et al.
Pubblicazione: (2024)
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
di: Wang, Junyu, et al.
Pubblicazione: (2025)
di: Wang, Junyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
di: Rong, Yan, et al.
Pubblicazione: (2025) -
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
di: Deng, Jiajun, et al.
Pubblicazione: (2025) -
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
di: Su, Fei, et al.
Pubblicazione: (2026) -
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
di: Lin, Yueqian, et al.
Pubblicazione: (2025) -
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
di: Deng, Zeyu, et al.
Pubblicazione: (2025)