TeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akkiraju, Bhavana, Bandarupalli, Srihari, Sambangi, Swathi, Ravuri, Vasavi, Saraswathi, R Vijaya, Vuppala, Anil Kumar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
von: Pothula, Aishwarya, et al.
Veröffentlicht: (2025)
von: Pothula, Aishwarya, et al.
Veröffentlicht: (2025)
Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data
von: Bandarupalli, Srihari, et al.
Veröffentlicht: (2025)
von: Bandarupalli, Srihari, et al.
Veröffentlicht: (2025)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus
von: Adetiba, Emmanuel, et al.
Veröffentlicht: (2025)
von: Adetiba, Emmanuel, et al.
Veröffentlicht: (2025)
Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
von: Akarsh, Sai, et al.
Veröffentlicht: (2024)
von: Akarsh, Sai, et al.
Veröffentlicht: (2024)
Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS
von: Anuprabha, M, et al.
Veröffentlicht: (2025)
von: Anuprabha, M, et al.
Veröffentlicht: (2025)
Typical vs. Atypical Disfluency Classification: Introducing the IIITH-TISA Corpus and Temporal Context-Based Feature Representations
von: Kommagouni, Priyanka, et al.
Veröffentlicht: (2024)
von: Kommagouni, Priyanka, et al.
Veröffentlicht: (2024)
A Multi-modal Approach to Dysarthria Detection and Severity Assessment Using Speech and Text Information
von: M, Anuprabha, et al.
Veröffentlicht: (2024)
von: M, Anuprabha, et al.
Veröffentlicht: (2024)
Open vocabulary keyword spotting through transfer learning from speech synthesis
von: V, Kesavaraj, et al.
Veröffentlicht: (2024)
von: V, Kesavaraj, et al.
Veröffentlicht: (2024)
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
von: Mondal, Anindita, et al.
Veröffentlicht: (2024)
von: Mondal, Anindita, et al.
Veröffentlicht: (2024)
End-to-End User-Defined Keyword Spotting using Shifted Delta Coefficients
von: V, Kesavaraj, et al.
Veröffentlicht: (2024)
von: V, Kesavaraj, et al.
Veröffentlicht: (2024)
Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech
von: Wotherspoon, Shannon, et al.
Veröffentlicht: (2024)
von: Wotherspoon, Shannon, et al.
Veröffentlicht: (2024)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
DiariST: Streaming Speech Translation with Speaker Diarization
von: Yang, Mu, et al.
Veröffentlicht: (2023)
von: Yang, Mu, et al.
Veröffentlicht: (2023)
Rethinking Cross-Corpus Speech Emotion Recognition Benchmarking: Are Paralinguistic Pre-Trained Representations Sufficient?
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
Enhancing Speaker-Independent Dysarthric Speech Severity Classification with DSSCNet and Cross-Corpus Adaptation
von: Roy, Arnab Kumar, et al.
Veröffentlicht: (2025)
von: Roy, Arnab Kumar, et al.
Veröffentlicht: (2025)
GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
UrduSpeech: A 156-Hour Urdu Speech Corpus with 12-Dimension Paralinguistic Annotations
von: Haq, Attia Nafees ul, et al.
Veröffentlicht: (2026)
von: Haq, Attia Nafees ul, et al.
Veröffentlicht: (2026)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
On Feature Learning for Titi Monkey Activity Detection
von: Ravuri, Aditya, et al.
Veröffentlicht: (2024)
von: Ravuri, Aditya, et al.
Veröffentlicht: (2024)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis
von: Take, Osamu, et al.
Veröffentlicht: (2024)
von: Take, Osamu, et al.
Veröffentlicht: (2024)
spINAch: A Diachronic Corpus of French Broadcast Speech Controlled for Speakers' Age and Gender
von: Devauchelle, Simon, et al.
Veröffentlicht: (2026)
von: Devauchelle, Simon, et al.
Veröffentlicht: (2026)
CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
von: Carvalho, Carlos, et al.
Veröffentlicht: (2025)
von: Carvalho, Carlos, et al.
Veröffentlicht: (2025)
Open-Source System for Multilingual Translation and Cloned Speech Synthesis
von: Cámara, Mateo, et al.
Veröffentlicht: (2025)
von: Cámara, Mateo, et al.
Veröffentlicht: (2025)
TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
Direct Speech-to-Speech Neural Machine Translation: A Survey
von: Gupta, Mahendra, et al.
Veröffentlicht: (2024)
von: Gupta, Mahendra, et al.
Veröffentlicht: (2024)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
von: Xin, Detai, et al.
Veröffentlicht: (2023)
von: Xin, Detai, et al.
Veröffentlicht: (2023)
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Benchmarking Neural Speech Codec Intelligibility with SITool
von: Leschanowsky, Anna, et al.
Veröffentlicht: (2025)
von: Leschanowsky, Anna, et al.
Veröffentlicht: (2025)
Enhancing ASR Performance in the Medical Domain for Dravidian Languages
von: Devarakonda, Sri Charan, et al.
Veröffentlicht: (2026)
von: Devarakonda, Sri Charan, et al.
Veröffentlicht: (2026)
ÌròyìnSpeech: A multi-purpose Yorùbá Speech Corpus
von: Ogunremi, Tolulope, et al.
Veröffentlicht: (2023)
von: Ogunremi, Tolulope, et al.
Veröffentlicht: (2023)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
von: Goswami, Mandip
Veröffentlicht: (2026)
von: Goswami, Mandip
Veröffentlicht: (2026)
Ähnliche Einträge
-
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
von: Pothula, Aishwarya, et al.
Veröffentlicht: (2025) -
Efficient ASR for Low-Resource Languages: Leveraging Cross-Lingual Unlabeled Data
von: Bandarupalli, Srihari, et al.
Veröffentlicht: (2025) -
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025) -
BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus
von: Adetiba, Emmanuel, et al.
Veröffentlicht: (2025) -
Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
von: Akarsh, Sai, et al.
Veröffentlicht: (2024)