M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen-Phuoc, Long, Gaboriau, Renald, Delacroix, Dimitri, Navarro, Laurent |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PTSD-MDNN : Fusion tardive de réseaux de neurones profonds multimodaux pour la détection du trouble de stress post-traumatique
von: Nguyen-Phuoc, Long, et al.
Veröffentlicht: (2024)
von: Nguyen-Phuoc, Long, et al.
Veröffentlicht: (2024)
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
von: Hong, Joanna, et al.
Veröffentlicht: (2025)
von: Hong, Joanna, et al.
Veröffentlicht: (2025)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
von: Xu, Le, et al.
Veröffentlicht: (2025)
von: Xu, Le, et al.
Veröffentlicht: (2025)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
von: Gu, Ke, et al.
Veröffentlicht: (2025)
von: Gu, Ke, et al.
Veröffentlicht: (2025)
AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
von: Cui, Meng, et al.
Veröffentlicht: (2023)
von: Cui, Meng, et al.
Veröffentlicht: (2023)
Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
von: Kim, Minyoung, et al.
Veröffentlicht: (2025)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation
von: Wang, Baisen, et al.
Veröffentlicht: (2024)
von: Wang, Baisen, et al.
Veröffentlicht: (2024)
MCDubber: Multimodal Context-Aware Expressive Video Dubbing
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
von: Zhao, Yuan, et al.
Veröffentlicht: (2024)
Video-Guided Foley Sound Generation with Multimodal Controls
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
Digit Recognition using Multimodal Spiking Neural Networks
von: Bjorndahl, William, et al.
Veröffentlicht: (2024)
von: Bjorndahl, William, et al.
Veröffentlicht: (2024)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
von: Yu, Fei, et al.
Veröffentlicht: (2024)
von: Yu, Fei, et al.
Veröffentlicht: (2024)
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
von: You, Wenhao, et al.
Veröffentlicht: (2025)
von: You, Wenhao, et al.
Veröffentlicht: (2025)
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
von: Zhao, Fuzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Fuzheng, et al.
Veröffentlicht: (2024)
Synchformer: Efficient Synchronization from Sparse Cues
von: Iashin, Vladimir, et al.
Veröffentlicht: (2024)
von: Iashin, Vladimir, et al.
Veröffentlicht: (2024)
Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation
von: Kim, Gwanghyun, et al.
Veröffentlicht: (2024)
von: Kim, Gwanghyun, et al.
Veröffentlicht: (2024)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
von: He, Yuhang, et al.
Veröffentlicht: (2024)
von: He, Yuhang, et al.
Veröffentlicht: (2024)
Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning
von: Singh, Nikhil, et al.
Veröffentlicht: (2023)
von: Singh, Nikhil, et al.
Veröffentlicht: (2023)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
von: Wang, Anna, et al.
Veröffentlicht: (2024)
von: Wang, Anna, et al.
Veröffentlicht: (2024)
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
von: Pham, Kien T., et al.
Veröffentlicht: (2025)
von: Pham, Kien T., et al.
Veröffentlicht: (2025)
CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization
von: Huang, Liangbin, et al.
Veröffentlicht: (2026)
von: Huang, Liangbin, et al.
Veröffentlicht: (2026)
A Survey on Multimodal Music Emotion Recognition
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
von: Liyanarachchi, Rashini, et al.
Veröffentlicht: (2025)
Personality-Enhanced Multimodal Depression Detection in the Elderly
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
von: Wang, Honghong, et al.
Veröffentlicht: (2025)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PTSD-MDNN : Fusion tardive de réseaux de neurones profonds multimodaux pour la détection du trouble de stress post-traumatique
von: Nguyen-Phuoc, Long, et al.
Veröffentlicht: (2024) -
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
von: Hong, Joanna, et al.
Veröffentlicht: (2025) -
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
von: Xu, Le, et al.
Veröffentlicht: (2025) -
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
von: Gu, Ke, et al.
Veröffentlicht: (2025) -
AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)