Gespeichert in:
| Hauptverfasser: | Zheng, Xianrui, Sun, Guangzhi, Zhang, Chao, Woodland, Philip C. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2407.02007 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DNCASR: End-to-End Training for Speaker-Attributed ASR
von: Zheng, Xianrui, et al.
Veröffentlicht: (2025)
von: Zheng, Xianrui, et al.
Veröffentlicht: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
von: Fan, Zhiyun, et al.
Veröffentlicht: (2024)
von: Fan, Zhiyun, et al.
Veröffentlicht: (2024)
Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression
von: Wu, Wen, et al.
Veröffentlicht: (2023)
von: Wu, Wen, et al.
Veröffentlicht: (2023)
Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
von: Lashkarashvili, Nineli, et al.
Veröffentlicht: (2024)
von: Lashkarashvili, Nineli, et al.
Veröffentlicht: (2024)
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
von: Deng, Keqi, et al.
Veröffentlicht: (2023)
von: Deng, Keqi, et al.
Veröffentlicht: (2023)
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
von: Morrone, Giovanni, et al.
Veröffentlicht: (2024)
von: Morrone, Giovanni, et al.
Veröffentlicht: (2024)
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
OCR-Enhanced Multimodal ASR Can Read While Listening
von: Chen, Junli, et al.
Veröffentlicht: (2026)
von: Chen, Junli, et al.
Veröffentlicht: (2026)
Multiplexing Neural Audio Watermarks
von: Yuan, Zheqi, et al.
Veröffentlicht: (2025)
von: Yuan, Zheqi, et al.
Veröffentlicht: (2025)
Distribution-based Emotion Recognition in Conversation
von: Wu, Wen, et al.
Veröffentlicht: (2022)
von: Wu, Wen, et al.
Veröffentlicht: (2022)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
von: Shao, Yiwen, et al.
Veröffentlicht: (2024)
von: Shao, Yiwen, et al.
Veröffentlicht: (2024)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
On Speaker Attribution with SURT
von: Raj, Desh, et al.
Veröffentlicht: (2024)
von: Raj, Desh, et al.
Veröffentlicht: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
Joint ASR and Speaker Role Tagging with Serialized Output Training
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
von: He, Xiluo, et al.
Veröffentlicht: (2025)
von: He, Xiluo, et al.
Veröffentlicht: (2025)
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2024)
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2024)
Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation
von: Cai, Danwei, et al.
Veröffentlicht: (2023)
von: Cai, Danwei, et al.
Veröffentlicht: (2023)
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
von: Polok, Alexander, et al.
Veröffentlicht: (2026)
Neural Forward Filtering for Speaker-Image Separation
von: Sun, Jingqi, et al.
Veröffentlicht: (2025)
von: Sun, Jingqi, et al.
Veröffentlicht: (2025)
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
von: Pražák, Aleš, et al.
Veröffentlicht: (2025)
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DNCASR: End-to-End Training for Speaker-Attributed ASR
von: Zheng, Xianrui, et al.
Veröffentlicht: (2025) -
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
von: Sun, Guangzhi, et al.
Veröffentlicht: (2022) -
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024) -
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
von: Wang, Mengqi, et al.
Veröffentlicht: (2025) -
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)