End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ghane, Mohsen, Safari, Mohammad Sadegh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
von: You, Zhenghai, et al.
Veröffentlicht: (2025)
Unmixing the Crowd: Learning Mixture-to-Set Speaker Embeddings for Enrollment-Free Target Speech Extraction
von: Sidharth, FNU, et al.
Veröffentlicht: (2026)
von: Sidharth, FNU, et al.
Veröffentlicht: (2026)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024)
von: Lin, Wan, et al.
Veröffentlicht: (2024)
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
von: Jin, Zhan, et al.
Veröffentlicht: (2025)
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
von: Zhuang, Zhuoran, et al.
Veröffentlicht: (2026)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
von: Zhang, Lin, et al.
Veröffentlicht: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions
von: Xue, Ke, et al.
Veröffentlicht: (2025)
von: Xue, Ke, et al.
Veröffentlicht: (2025)
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
From Modular to End-to-End Speaker Diarization
von: Landini, Federico
Veröffentlicht: (2024)
von: Landini, Federico
Veröffentlicht: (2024)
Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization
von: Singh, Prachi, et al.
Veröffentlicht: (2024)
von: Singh, Prachi, et al.
Veröffentlicht: (2024)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
von: You, Zhenghai, et al.
Veröffentlicht: (2025) -
Unmixing the Crowd: Learning Mixture-to-Set Speaker Embeddings for Enrollment-Free Target Speech Extraction
von: Sidharth, FNU, et al.
Veröffentlicht: (2026) -
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024) -
Robust Audio-Visual Target Speaker Extraction with Emotion-Aware Multiple Enrollment Fusion
von: Jin, Zhan, et al.
Veröffentlicht: (2025) -
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)