Using Phonemes in cascaded S2S translation pipeline
Fuente:
arXiv
Salvato in:
| Autori principali: | Pilz, Rene, Schneider, Johannes |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
di: Karapiperis, Sotirios, et al.
Pubblicazione: (2024)
di: Karapiperis, Sotirios, et al.
Pubblicazione: (2024)
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
di: Weise, Tobias, et al.
Pubblicazione: (2024)
di: Weise, Tobias, et al.
Pubblicazione: (2024)
Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings
di: Nallaguntla, Vamshi, et al.
Pubblicazione: (2026)
di: Nallaguntla, Vamshi, et al.
Pubblicazione: (2026)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
di: Yeo, Eunjung, et al.
Pubblicazione: (2026)
di: Yeo, Eunjung, et al.
Pubblicazione: (2026)
Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech
di: Kolani, Yakov, et al.
Pubblicazione: (2025)
di: Kolani, Yakov, et al.
Pubblicazione: (2025)
Multitask Learning for Grapheme-to-Phoneme Conversion of Anglicisms in German Speech Recognition
di: Pritzen, Julia, et al.
Pubblicazione: (2021)
di: Pritzen, Julia, et al.
Pubblicazione: (2021)
DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
di: Lee, Wonjun, et al.
Pubblicazione: (2025)
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
di: Ohnaka, Hien, et al.
Pubblicazione: (2025)
di: Ohnaka, Hien, et al.
Pubblicazione: (2025)
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
di: Poli, Maxime, et al.
Pubblicazione: (2026)
di: Poli, Maxime, et al.
Pubblicazione: (2026)
Phoneme Hallucinator: One-shot Voice Conversion via Set Expansion
di: Shan, Siyuan, et al.
Pubblicazione: (2023)
di: Shan, Siyuan, et al.
Pubblicazione: (2023)
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
di: Poli, Maxime, et al.
Pubblicazione: (2024)
di: Poli, Maxime, et al.
Pubblicazione: (2024)
The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset
di: Landau, Gilad, et al.
Pubblicazione: (2025)
di: Landau, Gilad, et al.
Pubblicazione: (2025)
Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
di: Nachmani, Eliya, et al.
Pubblicazione: (2023)
di: Nachmani, Eliya, et al.
Pubblicazione: (2023)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
di: Kim, Eungbeom, et al.
Pubblicazione: (2024)
di: Kim, Eungbeom, et al.
Pubblicazione: (2024)
Robust and Explainable Depression Identification from Speech Using Vowel-Based Ensemble Learning Approaches
di: Feng, Kexin, et al.
Pubblicazione: (2024)
di: Feng, Kexin, et al.
Pubblicazione: (2024)
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
di: Gállego, Gerard I., et al.
Pubblicazione: (2025)
di: Gállego, Gerard I., et al.
Pubblicazione: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
di: Papi, Sara, et al.
Pubblicazione: (2023)
di: Papi, Sara, et al.
Pubblicazione: (2023)
An approach to optimize inference of the DIART speaker diarization pipeline
di: Aperdannier, Roman, et al.
Pubblicazione: (2024)
di: Aperdannier, Roman, et al.
Pubblicazione: (2024)
Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS
di: Fetrat, Mahta, et al.
Pubblicazione: (2025)
di: Fetrat, Mahta, et al.
Pubblicazione: (2025)
Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
di: Dong, Lukuan, et al.
Pubblicazione: (2024)
di: Dong, Lukuan, et al.
Pubblicazione: (2024)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
Unified Learnable 2D Convolutional Feature Extraction for ASR
di: Vieting, Peter, et al.
Pubblicazione: (2025)
di: Vieting, Peter, et al.
Pubblicazione: (2025)
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
di: Kögel, Fabian, et al.
Pubblicazione: (2023)
di: Kögel, Fabian, et al.
Pubblicazione: (2023)
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
di: Ghosh, Sreyan, et al.
Pubblicazione: (2025)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2025)
The ART of Conversation: Measuring Phonetic Convergence and Deliberate Imitation in L2-Speech with a Siamese RNN
di: Yuan, Zheng, et al.
Pubblicazione: (2023)
di: Yuan, Zheng, et al.
Pubblicazione: (2023)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
di: Ngo, Huong, et al.
Pubblicazione: (2025)
di: Ngo, Huong, et al.
Pubblicazione: (2025)
Bayesian Low-Rank Factorization for Robust Model Adaptation
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
How Does a Deep Neural Network Look at Lexical Stress in English Words?
di: Allouche, Itai, et al.
Pubblicazione: (2025)
di: Allouche, Itai, et al.
Pubblicazione: (2025)
Voice Impression Control in Zero-Shot TTS
di: Fujita, Kenichi, et al.
Pubblicazione: (2025)
di: Fujita, Kenichi, et al.
Pubblicazione: (2025)
WhisperRT -- Turning Whisper into a Causal Streaming Model
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
di: Carbonneau, Marc-André, et al.
Pubblicazione: (2025)
di: Carbonneau, Marc-André, et al.
Pubblicazione: (2025)
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
di: Amooie, Reihaneh, et al.
Pubblicazione: (2025)
di: Amooie, Reihaneh, et al.
Pubblicazione: (2025)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
di: Eren, Eray, et al.
Pubblicazione: (2025)
di: Eren, Eray, et al.
Pubblicazione: (2025)
Label-Context-Dependent Internal Language Model Estimation for CTC
di: Yang, Zijian, et al.
Pubblicazione: (2025)
di: Yang, Zijian, et al.
Pubblicazione: (2025)
TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2025)
Error Analysis in a Modular Meeting Transcription System
di: Vieting, Peter, et al.
Pubblicazione: (2025)
di: Vieting, Peter, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
di: Fujita, Kenichi, et al.
Pubblicazione: (2024) -
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
di: Karapiperis, Sotirios, et al.
Pubblicazione: (2024) -
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
di: Weise, Tobias, et al.
Pubblicazione: (2024) -
Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings
di: Nallaguntla, Vamshi, et al.
Pubblicazione: (2026) -
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
di: Yeo, Eunjung, et al.
Pubblicazione: (2026)