Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Wonjun, Kim, San, Lee, Gary Geunbae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
An Investigation Into Explainable Audio Hate Speech Detection
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
von: An, Jinmyeong, et al.
Veröffentlicht: (2024)
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
von: Do, Heejin, et al.
Veröffentlicht: (2024)
von: Do, Heejin, et al.
Veröffentlicht: (2024)
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
von: Kim, Youngjae, et al.
Veröffentlicht: (2024)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
von: Wei, Kun, et al.
Veröffentlicht: (2023)
von: Wei, Kun, et al.
Veröffentlicht: (2023)
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
von: Shi, Xiaohan, et al.
Veröffentlicht: (2023)
Exploration of Adapter for Noise Robust Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
von: Lin, Zhennan, et al.
Veröffentlicht: (2025)
von: Lin, Zhennan, et al.
Veröffentlicht: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
von: Paik, Gio, et al.
Veröffentlicht: (2025)
von: Paik, Gio, et al.
Veröffentlicht: (2025)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data
von: Kim, Seung-Bin, et al.
Veröffentlicht: (2024)
von: Kim, Seung-Bin, et al.
Veröffentlicht: (2024)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
von: Zhang, Enshi, et al.
Veröffentlicht: (2024)
von: Zhang, Enshi, et al.
Veröffentlicht: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
von: Wang, He, et al.
Veröffentlicht: (2025)
von: Wang, He, et al.
Veröffentlicht: (2025)
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
von: Cha, Jun-Hyeok, et al.
Veröffentlicht: (2025)
von: Cha, Jun-Hyeok, et al.
Veröffentlicht: (2025)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Speaker-Aware Simulation Improves Conversational Speech Recognition
von: Gedeon, Máté, et al.
Veröffentlicht: (2026)
von: Gedeon, Máté, et al.
Veröffentlicht: (2026)
Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
von: Roll, Nathan, et al.
Veröffentlicht: (2025)
von: Roll, Nathan, et al.
Veröffentlicht: (2025)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
von: Wang, Shiyao, et al.
Veröffentlicht: (2024)
von: Wang, Shiyao, et al.
Veröffentlicht: (2024)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
von: He, Haorui, et al.
Veröffentlicht: (2024)
von: He, Haorui, et al.
Veröffentlicht: (2024)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2025) -
An Investigation Into Explainable Audio Hate Speech Detection
von: An, Jinmyeong, et al.
Veröffentlicht: (2024) -
Acoustic Feature Mixup for Balanced Multi-aspect Pronunciation Assessment
von: Do, Heejin, et al.
Veröffentlicht: (2024) -
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
von: Jeon, Yejin, et al.
Veröffentlicht: (2024) -
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)