Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pražák, Aleš, Kunešová, Marie, Psutka, Josef |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
von: Kunešová, Marie, et al.
Veröffentlicht: (2025)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
von: He, Xiluo, et al.
Veröffentlicht: (2025)
von: He, Xiluo, et al.
Veröffentlicht: (2025)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2025)
von: Ma, Hao, et al.
Veröffentlicht: (2025)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
von: Medennikov, Ivan, et al.
Veröffentlicht: (2025)
von: Medennikov, Ivan, et al.
Veröffentlicht: (2025)
Joint ASR and Speaker Role Tagging with Serialized Output Training
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions
von: Xue, Ke, et al.
Veröffentlicht: (2025)
von: Xue, Ke, et al.
Veröffentlicht: (2025)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
von: Fan, Zhiyun, et al.
Veröffentlicht: (2024)
von: Fan, Zhiyun, et al.
Veröffentlicht: (2024)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Bengali-Loop: Community Benchmarks for Long-Form Bangla ASR and Speaker Diarization
von: Tabib, H. M. Shadman, et al.
Veröffentlicht: (2026)
von: Tabib, H. M. Shadman, et al.
Veröffentlicht: (2026)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling
von: Wu, Shu, et al.
Veröffentlicht: (2025)
von: Wu, Shu, et al.
Veröffentlicht: (2025)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
von: Sato, Hiroshi, et al.
Veröffentlicht: (2024)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2024)
AuralNet: Hierarchical Attention-based 3D Binaural Localization of Overlapping Speakers
von: Fu, Linya, et al.
Veröffentlicht: (2025)
von: Fu, Linya, et al.
Veröffentlicht: (2025)
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
von: Zeng, Bang, et al.
Veröffentlicht: (2024)
Target Speaker Extraction with Curriculum Learning
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2023)
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2023)
Towards Lightweight Speaker Verification via Adaptive Neural Network Quantization
von: Liu, Bei, et al.
Veröffentlicht: (2024)
von: Liu, Bei, et al.
Veröffentlicht: (2024)
Universal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
von: Zeng, Bang, et al.
Veröffentlicht: (2025)
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
Robust Target Speaker Direction of Arrival Estimation
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
von: Li, Longhao, et al.
Veröffentlicht: (2025)
von: Li, Longhao, et al.
Veröffentlicht: (2025)
Technical Report: A Practical Guide to Kaldi ASR Optimization
von: Hong, Mengze, et al.
Veröffentlicht: (2025)
von: Hong, Mengze, et al.
Veröffentlicht: (2025)
Listen to Extract: Onset-Prompted Target Speaker Extraction
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
Binaural Target Speaker Extraction using Individualized HRTF
von: Ellinson, Yoav, et al.
Veröffentlicht: (2025)
von: Ellinson, Yoav, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
von: Kunešová, Marie, et al.
Veröffentlicht: (2025) -
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
von: Kunešová, Marie, et al.
Veröffentlicht: (2025) -
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024) -
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2025) -
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
von: He, Xiluo, et al.
Veröffentlicht: (2025)