Language translation, and change of accent for speech-to-speech task using diffusion model
Fuente:
arXiv
Saved in:
| Main Authors: | Mishra, Abhishek, Chowdhury, Ritesh Sur, Bahuguna, Vartul, Pandey, Isha, Ramakrishnan, Ganesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition
by: Pandey, Isha, et al.
Published: (2026)
by: Pandey, Isha, et al.
Published: (2026)
A2TTS: TTS for Low Resource Indian Languages
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
Acoustic and perceptual differences between standard and accented speech and their voice clones
by: Yang, Tianle, et al.
Published: (2026)
by: Yang, Tianle, et al.
Published: (2026)
Addressing speaker gender bias in large scale speech translation systems
by: Bansal, Shubham, et al.
Published: (2025)
by: Bansal, Shubham, et al.
Published: (2025)
Part-of-speech tagging for Nagamese Language using CRF
by: Shohe, Alovi N, et al.
Published: (2025)
by: Shohe, Alovi N, et al.
Published: (2025)
The order in speech disorder: a scoping review of state of the art machine learning methods for clinical speech classification
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
by: Zhao, Rui, et al.
Published: (2024)
by: Zhao, Rui, et al.
Published: (2024)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
by: Jung, Jee-weon, et al.
Published: (2024)
by: Jung, Jee-weon, et al.
Published: (2024)
A cross-species neural foundation model for end-to-end speech decoding
by: Zhang, Yizi, et al.
Published: (2025)
by: Zhang, Yizi, et al.
Published: (2025)
Word stress in self-supervised speech models: A cross-linguistic comparison
by: Bentum, Martijn, et al.
Published: (2025)
by: Bentum, Martijn, et al.
Published: (2025)
Tagarela - A Portuguese speech dataset from podcasts
by: de Oliveira, Frederico Santos, et al.
Published: (2026)
by: de Oliveira, Frederico Santos, et al.
Published: (2026)
Employing self-supervised learning models for cross-linguistic child speech maturity classification
by: Zhang, Theo, et al.
Published: (2025)
by: Zhang, Theo, et al.
Published: (2025)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
by: Lonergan, Liam, et al.
Published: (2024)
by: Lonergan, Liam, et al.
Published: (2024)
How good is GPT at writing political speeches for the White House?
by: Savoy, Jacques
Published: (2024)
by: Savoy, Jacques
Published: (2024)
Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
VorTEX: Various overlap ratio for Target speech EXtraction
by: Oh, Ro-hoon, et al.
Published: (2026)
by: Oh, Ro-hoon, et al.
Published: (2026)
Forensic deepfake audio detection using segmental speech features
by: Yang, Tianle, et al.
Published: (2025)
by: Yang, Tianle, et al.
Published: (2025)
Decoding the decoder: Contextual sequence-to-sequence modeling for intracortical speech decoding
by: Olak, Michal, et al.
Published: (2026)
by: Olak, Michal, et al.
Published: (2026)
Tag and correct: high precision post-editing approach to correction of speech recognition errors
by: Ziętkiewicz, Tomasz
Published: (2024)
by: Ziętkiewicz, Tomasz
Published: (2024)
Tracking the emergence of linguistic structure in self-supervised models learning from speech
by: Kloots, Marianne de Heer, et al.
Published: (2026)
by: Kloots, Marianne de Heer, et al.
Published: (2026)
A Code Comprehension Benchmark for Large Language Models for Code
by: Havare, Jayant, et al.
Published: (2025)
by: Havare, Jayant, et al.
Published: (2025)
BBPE16: UTF-16-based byte-level byte-pair encoding for improved multilingual speech recognition
by: Kim, Hyunsik, et al.
Published: (2026)
by: Kim, Hyunsik, et al.
Published: (2026)
Classification is a RAG problem: A case study on hate speech detection
by: Willats, Richard, et al.
Published: (2025)
by: Willats, Richard, et al.
Published: (2025)
Development and multi-center evaluation of domain-adapted speech recognition for human-AI teaming in real-world gastrointestinal endoscopy
by: Yang, Ruijie, et al.
Published: (2026)
by: Yang, Ruijie, et al.
Published: (2026)
From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning
by: Neema, Nishit, et al.
Published: (2025)
by: Neema, Nishit, et al.
Published: (2025)
Can we trust AI to detect healthy multilingual English speakers among the cognitively impaired cohort in the UK? An investigation using real-world conversational speech
by: Pahar, Madhurananda, et al.
Published: (2026)
by: Pahar, Madhurananda, et al.
Published: (2026)
Enhancing nonnative speech perception and production through an AI-powered application
by: Georgiou, Georgios P.
Published: (2025)
by: Georgiou, Georgios P.
Published: (2025)
Moshi: a speech-text foundation model for real-time dialogue
by: Défossez, Alexandre, et al.
Published: (2024)
by: Défossez, Alexandre, et al.
Published: (2024)
What does it take to get state of the art in simultaneous speech-to-speech translation?
by: Wilmet, Vincent, et al.
Published: (2024)
by: Wilmet, Vincent, et al.
Published: (2024)
Investigating the impact of 2D gesture representation on co-speech gesture generation
by: Guichoux, Teo, et al.
Published: (2024)
by: Guichoux, Teo, et al.
Published: (2024)
NeuroVoz: a Castillian Spanish corpus of parkinsonian speech
by: Mendes-Laureano, Janaína, et al.
Published: (2024)
by: Mendes-Laureano, Janaína, et al.
Published: (2024)
Assessing and Mitigating Data Memorization Risks in Fine-Tuned Large Language Models
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
Do self-supervised speech and language models extract similar representations as human brain?
by: Chen, Peili, et al.
Published: (2023)
by: Chen, Peili, et al.
Published: (2023)
Exploring the traditional NMT model and Large Language Model for chat translation
by: Yang, Jinlong, et al.
Published: (2024)
by: Yang, Jinlong, et al.
Published: (2024)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
by: C, Anandh, et al.
Published: (2025)
by: C, Anandh, et al.
Published: (2025)
InstructAudio: Unified speech and music generation with natural language instruction
by: Qiang, Chunyu, et al.
Published: (2025)
by: Qiang, Chunyu, et al.
Published: (2025)
A unified front-end framework for English text-to-speech synthesis
by: Ying, Zelin, et al.
Published: (2023)
by: Ying, Zelin, et al.
Published: (2023)
Persuasion Games using Large Language Models
by: Ramani, Ganesh Prasath, et al.
Published: (2024)
by: Ramani, Ganesh Prasath, et al.
Published: (2024)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
by: Ogun, Sewade, et al.
Published: (2024)
by: Ogun, Sewade, et al.
Published: (2024)
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm
by: Shashidhar, Sarvesh, et al.
Published: (2025)
by: Shashidhar, Sarvesh, et al.
Published: (2025)
Similar Items
-
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition
by: Pandey, Isha, et al.
Published: (2026) -
A2TTS: TTS for Low Resource Indian Languages
by: Bhadoriya, Ayush Singh, et al.
Published: (2025) -
Acoustic and perceptual differences between standard and accented speech and their voice clones
by: Yang, Tianle, et al.
Published: (2026) -
Addressing speaker gender bias in large scale speech translation systems
by: Bansal, Shubham, et al.
Published: (2025) -
Part-of-speech tagging for Nagamese Language using CRF
by: Shohe, Alovi N, et al.
Published: (2025)