On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Jinchuan, Peng, Yifan, Chen, William, Choi, Kwanghee, Livescu, Karen, Watanabe, Shinji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
Towards Robust Speech Representation Learning for Thousands of Languages
von: Chen, William, et al.
Veröffentlicht: (2024)
von: Chen, William, et al.
Veröffentlicht: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
OpusLM: A Family of Open Unified Speech Language Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
On-device Streaming Discrete Speech Units
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
On The Landscape of Spoken Language Models: A Comprehensive Survey
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
von: Chou, Ju-Chieh, et al.
Veröffentlicht: (2023)
von: Chou, Ju-Chieh, et al.
Veröffentlicht: (2023)
Self-Supervised Speech Representations are More Phonetic than Semantic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
YODAS: Youtube-Oriented Dataset for Audio and Speech
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
An Empirical Recipe for Universal Phone Recognition
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
von: Chang, Xuankai, et al.
Veröffentlicht: (2024)
von: Chang, Xuankai, et al.
Veröffentlicht: (2024)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
von: Wang, Shih-Heng, et al.
Veröffentlicht: (2026)
von: Wang, Shih-Heng, et al.
Veröffentlicht: (2026)
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
von: Yan, Brian, et al.
Veröffentlicht: (2025)
von: Yan, Brian, et al.
Veröffentlicht: (2025)
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis
von: Chen, Yushen, et al.
Veröffentlicht: (2026)
von: Chen, Yushen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
von: Peng, Yifan, et al.
Veröffentlicht: (2025) -
Towards Robust Speech Representation Learning for Thousands of Languages
von: Chen, William, et al.
Veröffentlicht: (2024) -
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024) -
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
von: Shi, Jiatong, et al.
Veröffentlicht: (2024) -
Evaluating Self-Supervised Speech Models via Text-Based LLMS
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)