Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System
Fuente:
arXiv
Guardado en:
| Autores principales: | Manohar, Vimal, Chen, Szu-Jui, Wang, Zhiqi, Fujita, Yusuke, Watanabe, Shinji, Khudanpur, Sanjeev |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HLTCOE JHU Submission to the Voice Privacy Challenge 2024
por: Xinyuan, Henry Li, et al.
Publicado: (2024)
por: Xinyuan, Henry Li, et al.
Publicado: (2024)
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
por: Yang, Mu, et al.
Publicado: (2025)
por: Yang, Mu, et al.
Publicado: (2025)
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
por: Huang, Ruizhe, et al.
Publicado: (2024)
por: Huang, Ruizhe, et al.
Publicado: (2024)
Modeling Overlapped Speech with Shuffles
por: Wiesner, Matthew, et al.
Publicado: (2026)
por: Wiesner, Matthew, et al.
Publicado: (2026)
Adversarial Attacks and Defenses for Speech Recognition Systems
por: Żelasko, Piotr, et al.
Publicado: (2021)
por: Żelasko, Piotr, et al.
Publicado: (2021)
End-to-End Speech Recognition with Pre-trained Masked Language Model
por: Higuchi, Yosuke, et al.
Publicado: (2024)
por: Higuchi, Yosuke, et al.
Publicado: (2024)
Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
por: Lin, Guan-Ting, et al.
Publicado: (2025)
por: Lin, Guan-Ting, et al.
Publicado: (2025)
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
por: Fujita, Yuya, et al.
Publicado: (2024)
por: Fujita, Yuya, et al.
Publicado: (2024)
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
por: Hussein, Amir, et al.
Publicado: (2025)
por: Hussein, Amir, et al.
Publicado: (2025)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
por: Liu, Hexin, et al.
Publicado: (2024)
por: Liu, Hexin, et al.
Publicado: (2024)
Neural Blind Source Separation and Diarization for Distant Speech Recognition
por: Bando, Yoshiaki, et al.
Publicado: (2024)
por: Bando, Yoshiaki, et al.
Publicado: (2024)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
por: Su, Bo-Hao, et al.
Publicado: (2025)
por: Su, Bo-Hao, et al.
Publicado: (2025)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
por: Maekaku, Takashi, et al.
Publicado: (2025)
por: Maekaku, Takashi, et al.
Publicado: (2025)
Unsupervised Speech Enhancement using Data-defined Priors
por: Klement, Dominik, et al.
Publicado: (2025)
por: Klement, Dominik, et al.
Publicado: (2025)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
por: Tsunoo, Emiru, et al.
Publicado: (2023)
por: Tsunoo, Emiru, et al.
Publicado: (2023)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
por: Chen, Szu-Jui, et al.
Publicado: (2026)
por: Chen, Szu-Jui, et al.
Publicado: (2026)
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
por: Cornell, Samuele, et al.
Publicado: (2024)
por: Cornell, Samuele, et al.
Publicado: (2024)
Leveraging Content and Acoustic Representations for Speech Emotion Recognition
por: Dutta, Soumya, et al.
Publicado: (2024)
por: Dutta, Soumya, et al.
Publicado: (2024)
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
por: Huang, Ruizhe, et al.
Publicado: (2024)
por: Huang, Ruizhe, et al.
Publicado: (2024)
Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts
por: Garg, Ashi, et al.
Publicado: (2025)
por: Garg, Ashi, et al.
Publicado: (2025)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
por: Cornell, Samuele, et al.
Publicado: (2024)
por: Cornell, Samuele, et al.
Publicado: (2024)
Universal Speech Content Factorization
por: Xinyuan, Henry Li, et al.
Publicado: (2026)
por: Xinyuan, Henry Li, et al.
Publicado: (2026)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
por: Tsunoo, Emiru, et al.
Publicado: (2024)
por: Tsunoo, Emiru, et al.
Publicado: (2024)
Interspeech 2025 URGENT Speech Enhancement Challenge
por: Saijo, Kohei, et al.
Publicado: (2025)
por: Saijo, Kohei, et al.
Publicado: (2025)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
por: Ku, Pin-Jui, et al.
Publicado: (2024)
por: Ku, Pin-Jui, et al.
Publicado: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
por: Wang, Shih-heng, et al.
Publicado: (2024)
por: Wang, Shih-heng, et al.
Publicado: (2024)
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
por: Garg, Ashi, et al.
Publicado: (2025)
por: Garg, Ashi, et al.
Publicado: (2025)
Multi-blank Transducers for Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2022)
por: Xu, Hainan, et al.
Publicado: (2022)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
por: Zhang, Lin, et al.
Publicado: (2026)
por: Zhang, Lin, et al.
Publicado: (2026)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
por: Tian, Jingguang, et al.
Publicado: (2024)
por: Tian, Jingguang, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization
por: Wu, Yihan, et al.
Publicado: (2024)
por: Wu, Yihan, et al.
Publicado: (2024)
Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
por: Arcos-Holzinger, Sandra, et al.
Publicado: (2026)
por: Arcos-Holzinger, Sandra, et al.
Publicado: (2026)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
Uni-VERSA: Versatile Speech Assessment with a Unified Network
por: Shi, Jiatong, et al.
Publicado: (2025)
por: Shi, Jiatong, et al.
Publicado: (2025)
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
por: Chang, Xuankai, et al.
Publicado: (2024)
por: Chang, Xuankai, et al.
Publicado: (2024)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2025)
por: Sudo, Yui, et al.
Publicado: (2025)
Universal Score-based Speech Enhancement with High Content Preservation
por: Scheibler, Robin, et al.
Publicado: (2024)
por: Scheibler, Robin, et al.
Publicado: (2024)
Ejemplares similares
-
HLTCOE JHU Submission to the Voice Privacy Challenge 2024
por: Xinyuan, Henry Li, et al.
Publicado: (2024) -
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
por: Yang, Mu, et al.
Publicado: (2025) -
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
por: Huang, Ruizhe, et al.
Publicado: (2024) -
Modeling Overlapped Speech with Shuffles
por: Wiesner, Matthew, et al.
Publicado: (2026) -
Adversarial Attacks and Defenses for Speech Recognition Systems
por: Żelasko, Piotr, et al.
Publicado: (2021)