SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
Fuente:
arXiv
Salvato in:
| Autori principali: | Grossman, Raymond, Park, Taejin, Dhawan, Kunal, Titus, Andrew, Zhi, Sophia, Shchadilova, Yulia, Wang, Weiqing, Balam, Jagadeesh, Ginsburg, Boris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
di: Medennikov, Ivan, et al.
Pubblicazione: (2025)
di: Medennikov, Ivan, et al.
Pubblicazione: (2025)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
di: Wang, Weiqing, et al.
Pubblicazione: (2024)
di: Wang, Weiqing, et al.
Pubblicazione: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
di: Wang, Weiqing, et al.
Pubblicazione: (2025)
di: Wang, Weiqing, et al.
Pubblicazione: (2025)
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
di: Huang, He, et al.
Pubblicazione: (2024)
di: Huang, He, et al.
Pubblicazione: (2024)
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
di: Wang, Jinhan, et al.
Pubblicazione: (2024)
di: Wang, Jinhan, et al.
Pubblicazione: (2024)
Flexible Multichannel Speech Enhancement for Noise-Robust Frontend
di: Jukić, Ante, et al.
Pubblicazione: (2024)
di: Jukić, Ante, et al.
Pubblicazione: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
di: Park, Taejin, et al.
Pubblicazione: (2024)
di: Park, Taejin, et al.
Pubblicazione: (2024)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
Schrödinger Bridge for Generative Speech Enhancement
di: Jukić, Ante, et al.
Pubblicazione: (2024)
di: Jukić, Ante, et al.
Pubblicazione: (2024)
Visual-based spatial audio generation system for multi-speaker environments
di: Liu, Xiaojing, et al.
Pubblicazione: (2025)
di: Liu, Xiaojing, et al.
Pubblicazione: (2025)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
di: Han, Runduo, et al.
Pubblicazione: (2024)
di: Han, Runduo, et al.
Pubblicazione: (2024)
Target speaker anonymization in multi-speaker recordings
di: Tomashenko, Natalia, et al.
Pubblicazione: (2025)
di: Tomashenko, Natalia, et al.
Pubblicazione: (2025)
Hierarchical speaker representation for target speaker extraction
di: He, Shulin, et al.
Pubblicazione: (2022)
di: He, Shulin, et al.
Pubblicazione: (2022)
Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition
di: Noroozi, Vahid, et al.
Pubblicazione: (2023)
di: Noroozi, Vahid, et al.
Pubblicazione: (2023)
Text adaptation for speaker verification with speaker-text factorized embeddings
di: Yang, Yexin, et al.
Pubblicazione: (2025)
di: Yang, Yexin, et al.
Pubblicazione: (2025)
Improving curriculum learning for target speaker extraction with synthetic speakers
di: Liu, Yun, et al.
Pubblicazione: (2024)
di: Liu, Yun, et al.
Pubblicazione: (2024)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
di: Hu, Ke, et al.
Pubblicazione: (2025)
di: Hu, Ke, et al.
Pubblicazione: (2025)
Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
di: Boeddeker, Christoph, et al.
Pubblicazione: (2024)
di: Boeddeker, Christoph, et al.
Pubblicazione: (2024)
Training and Inference Efficiency of Encoder-Decoder Speech Models
di: Żelasko, Piotr, et al.
Pubblicazione: (2025)
di: Żelasko, Piotr, et al.
Pubblicazione: (2025)
Triage knowledge distillation for speaker verification
di: Kim, Ju-ho, et al.
Pubblicazione: (2026)
di: Kim, Ju-ho, et al.
Pubblicazione: (2026)
Privacy-oriented manipulation of speaker representations
di: Teixeira, Francisco, et al.
Pubblicazione: (2023)
di: Teixeira, Francisco, et al.
Pubblicazione: (2023)
Curriculum learning for self-supervised speaker verification
di: Heo, Hee-Soo, et al.
Pubblicazione: (2022)
di: Heo, Hee-Soo, et al.
Pubblicazione: (2022)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
di: Huang, Ziling, et al.
Pubblicazione: (2025)
di: Huang, Ziling, et al.
Pubblicazione: (2025)
How phonemes contribute to deep speaker models?
di: Li, Pengqi, et al.
Pubblicazione: (2024)
di: Li, Pengqi, et al.
Pubblicazione: (2024)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
di: Ogun, Sewade, et al.
Pubblicazione: (2024)
di: Ogun, Sewade, et al.
Pubblicazione: (2024)
Investigation of perception inconsistency in speaker embedding for asynchronous voice anonymization
di: Wang, Rui, et al.
Pubblicazione: (2025)
di: Wang, Rui, et al.
Pubblicazione: (2025)
Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2024)
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2024)
The importance of spatial and spectral information in multiple speaker tracking
di: Beit-On, Hanan, et al.
Pubblicazione: (2024)
di: Beit-On, Hanan, et al.
Pubblicazione: (2024)
EasyEyes: Online hearing research using speakers calibrated by phones
di: Vican, Ivan, et al.
Pubblicazione: (2025)
di: Vican, Ivan, et al.
Pubblicazione: (2025)
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
End-to-end multi-channel speaker extraction and binaural speech synthesis
di: Chi, Cheng, et al.
Pubblicazione: (2024)
di: Chi, Cheng, et al.
Pubblicazione: (2024)
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
di: Sun, Chang, et al.
Pubblicazione: (2024)
di: Sun, Chang, et al.
Pubblicazione: (2024)
On the calibration of powerset speaker diarization models
di: Plaquet, Alexis, et al.
Pubblicazione: (2024)
di: Plaquet, Alexis, et al.
Pubblicazione: (2024)
A Benchmark for Multi-speaker Anonymization
di: Miao, Xiaoxiao, et al.
Pubblicazione: (2024)
di: Miao, Xiaoxiao, et al.
Pubblicazione: (2024)
On the influence of language similarity in non-target speaker verification trials
di: Reuter, Paul M., et al.
Pubblicazione: (2025)
di: Reuter, Paul M., et al.
Pubblicazione: (2025)
Audio-visual child-adult speaker classification in dyadic interactions
di: Xu, Anfeng, et al.
Pubblicazione: (2023)
di: Xu, Anfeng, et al.
Pubblicazione: (2023)
Spoken language change detection inspired by speaker change detection
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
Challenging margin-based speaker embedding extractors by using the variational information bottleneck
di: Stafylakis, Themos, et al.
Pubblicazione: (2024)
di: Stafylakis, Themos, et al.
Pubblicazione: (2024)
Tandem spoofing-robust automatic speaker verification based on time-domain embeddings
di: Weizman, Avishai, et al.
Pubblicazione: (2024)
di: Weizman, Avishai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
di: Medennikov, Ivan, et al.
Pubblicazione: (2025) -
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
di: Wang, Weiqing, et al.
Pubblicazione: (2024) -
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
di: Wang, Weiqing, et al.
Pubblicazione: (2025) -
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
di: Huang, He, et al.
Pubblicazione: (2024) -
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
di: Wang, Jinhan, et al.
Pubblicazione: (2024)