Enregistré dans:
| Auteurs principaux: | Khadse, Parth, Kopparapu, Sunil Kumar |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2602.14664 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Calculus-Based Framework for Determining Vocabulary Size in End-to-End ASR
par: Kopparapu, Sunil Kumar
Publié: (2026)
par: Kopparapu, Sunil Kumar
Publié: (2026)
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
par: Kopparapu, Sunil Kumar, et autres
Publié: (2024)
par: Kopparapu, Sunil Kumar, et autres
Publié: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
par: Guo, Yinlin, et autres
Publié: (2024)
par: Guo, Yinlin, et autres
Publié: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
par: Bataev, Vladimir, et autres
Publié: (2025)
par: Bataev, Vladimir, et autres
Publié: (2025)
Unifying EEG and Speech for Emotion Recognition: A Two-Step Joint Learning Framework for Handling Missing EEG Data During Inference
par: Tiwari, Upasana, et autres
Publié: (2025)
par: Tiwari, Upasana, et autres
Publié: (2025)
End-to-End Speech-to-Text Translation: A Survey
par: Sethiya, Nivedita, et autres
Publié: (2023)
par: Sethiya, Nivedita, et autres
Publié: (2023)
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
par: Tiwari, Upasana, et autres
Publié: (2025)
par: Tiwari, Upasana, et autres
Publié: (2025)
An investigation of phrase break prediction in an End-to-End TTS system
par: Vadapalli, Anandaswarup
Publié: (2023)
par: Vadapalli, Anandaswarup
Publié: (2023)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
par: Ahmad, Hawraz A., et autres
Publié: (2024)
par: Ahmad, Hawraz A., et autres
Publié: (2024)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
par: Chen, Junyang, et autres
Publié: (2026)
par: Chen, Junyang, et autres
Publié: (2026)
SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance
par: Lou, Haowei, et autres
Publié: (2025)
par: Lou, Haowei, et autres
Publié: (2025)
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
par: Min, Anna, et autres
Publié: (2025)
par: Min, Anna, et autres
Publié: (2025)
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
par: Wang, Jialing, et autres
Publié: (2026)
par: Wang, Jialing, et autres
Publié: (2026)
SAND Challenge: Four Approaches for Dysartria Severity Classification
par: Deshpande, Gauri, et autres
Publié: (2025)
par: Deshpande, Gauri, et autres
Publié: (2025)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
par: Huang, Wuwei, et autres
Publié: (2025)
par: Huang, Wuwei, et autres
Publié: (2025)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
par: Vendrame, Katia, et autres
Publié: (2025)
par: Vendrame, Katia, et autres
Publié: (2025)
Representation Purification for End-to-End Speech Translation
par: Zhang, Chengwei, et autres
Publié: (2024)
par: Zhang, Chengwei, et autres
Publié: (2024)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
par: Lu, Wenhuan, et autres
Publié: (2025)
par: Lu, Wenhuan, et autres
Publié: (2025)
Deep Speech Synthesis from Multimodal Articulatory Representations
par: Wu, Peter, et autres
Publié: (2024)
par: Wu, Peter, et autres
Publié: (2024)
Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics
par: Zhang, Ziqian, et autres
Publié: (2025)
par: Zhang, Ziqian, et autres
Publié: (2025)
Improved Dysarthric Speech to Text Conversion via TTS Personalization
par: Mihajlik, Péter, et autres
Publié: (2025)
par: Mihajlik, Péter, et autres
Publié: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
par: Liu, Huadai, et autres
Publié: (2023)
par: Liu, Huadai, et autres
Publié: (2023)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
par: Qi, Xin, et autres
Publié: (2024)
par: Qi, Xin, et autres
Publié: (2024)
DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
par: Yin, Kang, et autres
Publié: (2025)
par: Yin, Kang, et autres
Publié: (2025)
MunTTS: A Text-to-Speech System for Mundari
par: Gumma, Varun, et autres
Publié: (2024)
par: Gumma, Varun, et autres
Publié: (2024)
An End-to-End Speech Summarization Using Large Language Model
par: Shang, Hengchao, et autres
Publié: (2024)
par: Shang, Hengchao, et autres
Publié: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
par: Gupta, Kishan, et autres
Publié: (2024)
par: Gupta, Kishan, et autres
Publié: (2024)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
par: Zhou, Xuanru, et autres
Publié: (2024)
par: Zhou, Xuanru, et autres
Publié: (2024)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
par: Zhou, Junzuo, et autres
Publié: (2024)
par: Zhou, Junzuo, et autres
Publié: (2024)
Recent Advances in End-to-End Simultaneous Speech Translation
par: Liu, Xiaoqian, et autres
Publié: (2024)
par: Liu, Xiaoqian, et autres
Publié: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
par: Ren, Yong, et autres
Publié: (2026)
par: Ren, Yong, et autres
Publié: (2026)
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
par: Liang, Qifan, et autres
Publié: (2026)
par: Liang, Qifan, et autres
Publié: (2026)
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
par: Weise, Tobias, et autres
Publié: (2024)
par: Weise, Tobias, et autres
Publié: (2024)
Pushing the Limits of End-to-End Diarization
par: Broughton, Samuel J., et autres
Publié: (2025)
par: Broughton, Samuel J., et autres
Publié: (2025)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
par: Li, Tianpeng, et autres
Publié: (2025)
par: Li, Tianpeng, et autres
Publié: (2025)
Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
par: Raimon, Athul, et autres
Publié: (2024)
par: Raimon, Athul, et autres
Publié: (2024)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
par: Huang, Wuwei, et autres
Publié: (2025)
par: Huang, Wuwei, et autres
Publié: (2025)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
par: Bondaruk, Łukasz, et autres
Publié: (2025)
par: Bondaruk, Łukasz, et autres
Publié: (2025)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
par: Ma, Zhengrui, et autres
Publié: (2024)
par: Ma, Zhengrui, et autres
Publié: (2024)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
par: Guimarães, Heitor R., et autres
Publié: (2024)
par: Guimarães, Heitor R., et autres
Publié: (2024)
Documents similaires
-
A Calculus-Based Framework for Determining Vocabulary Size in End-to-End ASR
par: Kopparapu, Sunil Kumar
Publié: (2026) -
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
par: Kopparapu, Sunil Kumar, et autres
Publié: (2024) -
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
par: Guo, Yinlin, et autres
Publié: (2024) -
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
par: Bataev, Vladimir, et autres
Publié: (2025) -
Unifying EEG and Speech for Emotion Recognition: A Two-Step Joint Learning Framework for Handling Missing EEG Data During Inference
par: Tiwari, Upasana, et autres
Publié: (2025)