M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Wu, Shilong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
von: Morrone, Giovanni, et al.
Veröffentlicht: (2024)
von: Morrone, Giovanni, et al.
Veröffentlicht: (2024)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
Target Speech Diarization with Multimodal Prompts
von: Jiang, Yidi, et al.
Veröffentlicht: (2024)
von: Jiang, Yidi, et al.
Veröffentlicht: (2024)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
von: Mingote, Victoria, et al.
Veröffentlicht: (2024)
von: Mingote, Victoria, et al.
Veröffentlicht: (2024)
M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
pyAMPACT: A Score-Audio Alignment Toolkit for Performance Data Estimation and Multi-modal Processing
von: Devaney, Johanna, et al.
Veröffentlicht: (2024)
von: Devaney, Johanna, et al.
Veröffentlicht: (2024)
MMSD-Net: Towards Multi-modal Stuttering Detection
von: Nie, Liangyu, et al.
Veröffentlicht: (2024)
von: Nie, Liangyu, et al.
Veröffentlicht: (2024)
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
von: Wang, Anna, et al.
Veröffentlicht: (2024)
von: Wang, Anna, et al.
Veröffentlicht: (2024)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Attentive-based Multi-level Feature Fusion for Voice Disorder Diagnosis
von: Shen, Lipeng, et al.
Veröffentlicht: (2024)
von: Shen, Lipeng, et al.
Veröffentlicht: (2024)
Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification
von: Gu, Bin, et al.
Veröffentlicht: (2025)
von: Gu, Bin, et al.
Veröffentlicht: (2025)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
360+x: A Panoptic Multi-modal Scene Understanding Dataset
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Who is Authentic Speaker
von: Huang, Qiang
Veröffentlicht: (2024)
von: Huang, Qiang
Veröffentlicht: (2024)
Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
von: Cheng, Ming, et al.
Veröffentlicht: (2024)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
IML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech Processing
von: Song, Zeyang, et al.
Veröffentlicht: (2025)
von: Song, Zeyang, et al.
Veröffentlicht: (2025)
Jamendo-QA: A Large-Scale Music Question Answering Dataset
von: Koh, Junyoung, et al.
Veröffentlicht: (2025)
von: Koh, Junyoung, et al.
Veröffentlicht: (2025)
POLIPHONE: A Dataset for Smartphone Model Identification from Audio Recordings
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
von: Han, Jiangyu, et al.
Veröffentlicht: (2025)
CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework
von: Ning, Jinzhong, et al.
Veröffentlicht: (2025)
von: Ning, Jinzhong, et al.
Veröffentlicht: (2025)
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
von: Salganik, Rebecca, et al.
Veröffentlicht: (2026)
von: Salganik, Rebecca, et al.
Veröffentlicht: (2026)
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
von: Morrone, Giovanni, et al.
Veröffentlicht: (2024) -
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024) -
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2023) -
Target Speech Diarization with Multimodal Prompts
von: Jiang, Yidi, et al.
Veröffentlicht: (2024) -
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)