Transducer Consistency Regularization for Speech to Text Applications
Fuente:
arXiv
Salvato in:
| Autori principali: | Tseng, Cindy, Tang, Yun, Apsingekar, Vijendra Raj |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
di: Tang, Yun, et al.
Pubblicazione: (2025)
di: Tang, Yun, et al.
Pubblicazione: (2025)
Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
di: Tang, Yun, et al.
Pubblicazione: (2025)
di: Tang, Yun, et al.
Pubblicazione: (2025)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
di: Hussein, Amir, et al.
Pubblicazione: (2025)
di: Hussein, Amir, et al.
Pubblicazione: (2025)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
di: Andrusenko, Andrei, et al.
Pubblicazione: (2026)
di: Andrusenko, Andrei, et al.
Pubblicazione: (2026)
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
di: Zhang, Tian-Hao, et al.
Pubblicazione: (2023)
di: Zhang, Tian-Hao, et al.
Pubblicazione: (2023)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
di: Tseng, Yuan, et al.
Pubblicazione: (2025)
di: Tseng, Yuan, et al.
Pubblicazione: (2025)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2025)
Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
di: Allbert, Rumi, et al.
Pubblicazione: (2025)
di: Allbert, Rumi, et al.
Pubblicazione: (2025)
Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models
di: Tang, Zhiyuan, et al.
Pubblicazione: (2024)
di: Tang, Zhiyuan, et al.
Pubblicazione: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2024)
di: Xu, Hainan, et al.
Pubblicazione: (2024)
Revisiting Interpolation Augmentation for Speech-to-Text Generation
di: Xu, Chen, et al.
Pubblicazione: (2024)
di: Xu, Chen, et al.
Pubblicazione: (2024)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
Promptformer: Prompted Conformer Transducer for ASR
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
di: Wu, Haibin, et al.
Pubblicazione: (2025)
di: Wu, Haibin, et al.
Pubblicazione: (2025)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
di: Wang, Guansu, et al.
Pubblicazione: (2025)
di: Wang, Guansu, et al.
Pubblicazione: (2025)
Continuous Speech Tokenizer in Text To Speech
di: Li, Yixing, et al.
Pubblicazione: (2024)
di: Li, Yixing, et al.
Pubblicazione: (2024)
What Do Speech Foundation Models Not Learn About Speech?
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
Linguistic Knowledge Transfer Learning for Speech Enhancement
di: Hung, Kuo-Hsuan, et al.
Pubblicazione: (2025)
di: Hung, Kuo-Hsuan, et al.
Pubblicazione: (2025)
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
Lightweight Transducer Based on Frame-Level Criterion
di: Wan, Genshun, et al.
Pubblicazione: (2024)
di: Wan, Genshun, et al.
Pubblicazione: (2024)
SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis
di: Buess, Lukas, et al.
Pubblicazione: (2025)
di: Buess, Lukas, et al.
Pubblicazione: (2025)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025)
di: Futami, Hayato, et al.
Pubblicazione: (2025)
Estimating the Completeness of Discrete Speech Units
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2024)
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2024)
Self-Supervised Learning for Multi-Channel Neural Transducer
di: Kojima, Atsushi
Pubblicazione: (2024)
di: Kojima, Atsushi
Pubblicazione: (2024)
MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
di: Singh, Jaskaran, et al.
Pubblicazione: (2025)
di: Singh, Jaskaran, et al.
Pubblicazione: (2025)
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
di: Mondal, Anindita, et al.
Pubblicazione: (2024)
di: Mondal, Anindita, et al.
Pubblicazione: (2024)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
di: Baas, Matthew, et al.
Pubblicazione: (2025)
di: Baas, Matthew, et al.
Pubblicazione: (2025)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
di: Sudo, Yui, et al.
Pubblicazione: (2024)
di: Sudo, Yui, et al.
Pubblicazione: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
What do Speech Foundation Models Learn? Analysis and Applications
di: Pasad, Ankita
Pubblicazione: (2025)
di: Pasad, Ankita
Pubblicazione: (2025)
Documenti analoghi
-
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
di: Tang, Yun, et al.
Pubblicazione: (2025) -
Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
di: Tang, Yun, et al.
Pubblicazione: (2025) -
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
di: Deng, Keqi, et al.
Pubblicazione: (2024) -
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2024) -
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
di: Tseng, Wei-Cheng, et al.
Pubblicazione: (2025)