Multi-modal expressive personality recognition in data non-ideal audiovisual based on multi-scale feature enhancement and modal augment
Fuente:
arXiv
Salvato in:
| Autori principali: | Kong, Weixuan, Yu, Jinpeng, Li, Zijun, Liu, Hanwei, Qu, Jiqing, Xiao, Hui, Li, Xuefeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
di: Li, Xueyan, et al.
Pubblicazione: (2025)
di: Li, Xueyan, et al.
Pubblicazione: (2025)
FOCAL: A Novel Benchmarking Technique for Multi-modal Agents
di: Purwar, Anupam, et al.
Pubblicazione: (2026)
di: Purwar, Anupam, et al.
Pubblicazione: (2026)
Synaspot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy
di: Li, Kewei, et al.
Pubblicazione: (2025)
di: Li, Kewei, et al.
Pubblicazione: (2025)
DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
di: Yin, Kang, et al.
Pubblicazione: (2025)
di: Yin, Kang, et al.
Pubblicazione: (2025)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
di: Wang, Anna, et al.
Pubblicazione: (2024)
di: Wang, Anna, et al.
Pubblicazione: (2024)
Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction task
di: Coldenhoff, Jozef, et al.
Pubblicazione: (2024)
di: Coldenhoff, Jozef, et al.
Pubblicazione: (2024)
Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning
di: Pan, Xiaofeng, et al.
Pubblicazione: (2025)
di: Pan, Xiaofeng, et al.
Pubblicazione: (2025)
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
di: Wang, Xiaopeng, et al.
Pubblicazione: (2025)
di: Wang, Xiaopeng, et al.
Pubblicazione: (2025)
Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations
di: Fu, Jingwen, et al.
Pubblicazione: (2025)
di: Fu, Jingwen, et al.
Pubblicazione: (2025)
Multi-channel multi-speaker transformer for speech recognition
di: Yifan, Guo, et al.
Pubblicazione: (2026)
di: Yifan, Guo, et al.
Pubblicazione: (2026)
A multi-modal approach for identifying schizophrenia using cross-modal attention
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
Multi-modal Speech Enhancement with Limited Electromyography Channels
di: Feng, Fuyuan, et al.
Pubblicazione: (2025)
di: Feng, Fuyuan, et al.
Pubblicazione: (2025)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
di: Mu, Bingshen, et al.
Pubblicazione: (2024)
di: Mu, Bingshen, et al.
Pubblicazione: (2024)
Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation
di: Ai, Zhiqi, et al.
Pubblicazione: (2026)
di: Ai, Zhiqi, et al.
Pubblicazione: (2026)
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
di: Li, Junjie, et al.
Pubblicazione: (2024)
di: Li, Junjie, et al.
Pubblicazione: (2024)
Shared Multi-modal Embedding Space for Face-Voice Association
di: Simic, Christopher, et al.
Pubblicazione: (2025)
di: Simic, Christopher, et al.
Pubblicazione: (2025)
An efficient text augmentation approach for contextualized Mandarin speech recognition
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
Improving child speech recognition with augmented child-like speech
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
360+x: A Panoptic Multi-modal Scene Understanding Dataset
di: Chen, Hao, et al.
Pubblicazione: (2024)
di: Chen, Hao, et al.
Pubblicazione: (2024)
MMSD-Net: Towards Multi-modal Stuttering Detection
di: Nie, Liangyu, et al.
Pubblicazione: (2024)
di: Nie, Liangyu, et al.
Pubblicazione: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
di: Janiczek, John, et al.
Pubblicazione: (2024)
di: Janiczek, John, et al.
Pubblicazione: (2024)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
di: Li, Junjie, et al.
Pubblicazione: (2024)
di: Li, Junjie, et al.
Pubblicazione: (2024)
Cross-lingual Alzheimer's Disease detection based on paralinguistic and pre-trained features
di: Chen, Xuchu, et al.
Pubblicazione: (2023)
di: Chen, Xuchu, et al.
Pubblicazione: (2023)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
di: Zhou, Dingkun, et al.
Pubblicazione: (2025)
di: Zhou, Dingkun, et al.
Pubblicazione: (2025)
HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
di: Dutta, Soumya, et al.
Pubblicazione: (2023)
di: Dutta, Soumya, et al.
Pubblicazione: (2023)
Adversarial multi-task underwater acoustic target recognition: towards robustness against various influential factors
di: Xie, Yuan, et al.
Pubblicazione: (2024)
di: Xie, Yuan, et al.
Pubblicazione: (2024)
Towards interpretable emotion recognition: Identifying key features with machine learning
di: Kaloga, Yacouba, et al.
Pubblicazione: (2025)
di: Kaloga, Yacouba, et al.
Pubblicazione: (2025)
Mozart's Touch: A Lightweight Multi-modal Music Generation Framework Based on Pre-Trained Large Models
di: Li, Jiajun, et al.
Pubblicazione: (2024)
di: Li, Jiajun, et al.
Pubblicazione: (2024)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
di: Sadok, Samir, et al.
Pubblicazione: (2023)
di: Sadok, Samir, et al.
Pubblicazione: (2023)
An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
di: Inoue, Koji, et al.
Pubblicazione: (2025)
di: Inoue, Koji, et al.
Pubblicazione: (2025)
Knowledge-Decoupled Functionally Invariant Path with Synthetic Personal Data for Personalized ASR
di: Gu, Yue, et al.
Pubblicazione: (2025)
di: Gu, Yue, et al.
Pubblicazione: (2025)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
di: Tian, Zeyue, et al.
Pubblicazione: (2026)
di: Tian, Zeyue, et al.
Pubblicazione: (2026)
MM-KWS: Multi-modal Prompts for Multilingual User-defined Keyword Spotting
di: Ai, Zhiqi, et al.
Pubblicazione: (2024)
di: Ai, Zhiqi, et al.
Pubblicazione: (2024)
learning discriminative features from spectrograms using center loss for speech emotion recognition
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
di: Chuchra, Akanksha, et al.
Pubblicazione: (2026)
di: Chuchra, Akanksha, et al.
Pubblicazione: (2026)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
di: Liu, Shansong, et al.
Pubblicazione: (2023)
di: Liu, Shansong, et al.
Pubblicazione: (2023)
Robustifying automatic speech recognition by extracting slowly varying features
di: Pizarro, Matías, et al.
Pubblicazione: (2021)
di: Pizarro, Matías, et al.
Pubblicazione: (2021)
Documenti analoghi
-
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
di: Liu, Rui, et al.
Pubblicazione: (2024) -
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
di: Li, Xueyan, et al.
Pubblicazione: (2025) -
FOCAL: A Novel Benchmarking Technique for Multi-modal Agents
di: Purwar, Anupam, et al.
Pubblicazione: (2026) -
Synaspot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy
di: Li, Kewei, et al.
Pubblicazione: (2025) -
DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
di: Yin, Kang, et al.
Pubblicazione: (2025)