Target Speaker ASR with Whisper
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Polok, Alexander, Klement, Dominik, Wiesner, Matthew, Khudanpur, Sanjeev, Černocký, Jan, Burget, Lukáš |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
Unsupervised Speech Enhancement using Data-defined Priors
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Modeling Overlapped Speech with Shuffles
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
On Speaker Attribution with SURT
von: Raj, Desh, et al.
Veröffentlicht: (2024)
von: Raj, Desh, et al.
Veröffentlicht: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
BUT System for the MLC-SLM Challenge
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
von: He, Xiluo, et al.
Veröffentlicht: (2025)
von: He, Xiluo, et al.
Veröffentlicht: (2025)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR
von: Tan, Kai, et al.
Veröffentlicht: (2026)
von: Tan, Kai, et al.
Veröffentlicht: (2026)
Improving Automatic Speech Recognition with Decoder-Centric Regularisation in Encoder-Decoder Models
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
Leveraging Self-Supervised Learning for Speaker Diarization
von: Han, Jiangyu, et al.
Veröffentlicht: (2024)
von: Han, Jiangyu, et al.
Veröffentlicht: (2024)
SpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays
von: Shao, Yiwen, et al.
Veröffentlicht: (2026)
von: Shao, Yiwen, et al.
Veröffentlicht: (2026)
DeCRED: Decoder-Centric Regularization for Encoder-Decoder Based Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
von: Polok, Alexander, et al.
Veröffentlicht: (2025)
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
von: Zhang, Li, et al.
Veröffentlicht: (2024)
von: Zhang, Li, et al.
Veröffentlicht: (2024)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
BUT System Description for CHiME-9 MCoRec Challenge
von: Klement, Dominik, et al.
Veröffentlicht: (2026)
von: Klement, Dominik, et al.
Veröffentlicht: (2026)
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
Universal Speech Content Factorization
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2026)
von: Xinyuan, Henry Li, et al.
Veröffentlicht: (2026)
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2025)
von: Ma, Hao, et al.
Veröffentlicht: (2025)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
von: Garg, Ashi, et al.
Veröffentlicht: (2025)
von: Garg, Ashi, et al.
Veröffentlicht: (2025)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
von: Shao, Yiwen, et al.
Veröffentlicht: (2024)
von: Shao, Yiwen, et al.
Veröffentlicht: (2024)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Text-dependent Speaker Verification (TdSV) Challenge 2024: Challenge Evaluation Plan
von: Hossein, Zeinali, et al.
Veröffentlicht: (2024)
von: Hossein, Zeinali, et al.
Veröffentlicht: (2024)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Probing Self-supervised Learning Models with Target Speech Extraction
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors
von: Landini, Federico, et al.
Veröffentlicht: (2023)
von: Landini, Federico, et al.
Veröffentlicht: (2023)
Joint ASR and Speaker Role Tagging with Serialized Output Training
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
von: Udupa, Sathvik, et al.
Veröffentlicht: (2025)
von: Udupa, Sathvik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024) -
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2026) -
Unsupervised Speech Enhancement using Data-defined Priors
von: Klement, Dominik, et al.
Veröffentlicht: (2025) -
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025) -
Modeling Overlapped Speech with Shuffles
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)