Language translation, and change of accent for speech-to-speech task using diffusion model
Fuente:
arXiv
Guardado en:
| Autores principales: | Mishra, Abhishek, Chowdhury, Ritesh Sur, Bahuguna, Vartul, Pandey, Isha, Ramakrishnan, Ganesh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition
por: Pandey, Isha, et al.
Publicado: (2026)
por: Pandey, Isha, et al.
Publicado: (2026)
A2TTS: TTS for Low Resource Indian Languages
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025)
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025)
Acoustic and perceptual differences between standard and accented speech and their voice clones
por: Yang, Tianle, et al.
Publicado: (2026)
por: Yang, Tianle, et al.
Publicado: (2026)
Addressing speaker gender bias in large scale speech translation systems
por: Bansal, Shubham, et al.
Publicado: (2025)
por: Bansal, Shubham, et al.
Publicado: (2025)
Part-of-speech tagging for Nagamese Language using CRF
por: Shohe, Alovi N, et al.
Publicado: (2025)
por: Shohe, Alovi N, et al.
Publicado: (2025)
The order in speech disorder: a scoping review of state of the art machine learning methods for clinical speech classification
por: Moell, Birger, et al.
Publicado: (2025)
por: Moell, Birger, et al.
Publicado: (2025)
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
por: Zhao, Rui, et al.
Publicado: (2024)
por: Zhao, Rui, et al.
Publicado: (2024)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
por: Jung, Jee-weon, et al.
Publicado: (2024)
por: Jung, Jee-weon, et al.
Publicado: (2024)
A cross-species neural foundation model for end-to-end speech decoding
por: Zhang, Yizi, et al.
Publicado: (2025)
por: Zhang, Yizi, et al.
Publicado: (2025)
Word stress in self-supervised speech models: A cross-linguistic comparison
por: Bentum, Martijn, et al.
Publicado: (2025)
por: Bentum, Martijn, et al.
Publicado: (2025)
Tagarela - A Portuguese speech dataset from podcasts
por: de Oliveira, Frederico Santos, et al.
Publicado: (2026)
por: de Oliveira, Frederico Santos, et al.
Publicado: (2026)
Employing self-supervised learning models for cross-linguistic child speech maturity classification
por: Zhang, Theo, et al.
Publicado: (2025)
por: Zhang, Theo, et al.
Publicado: (2025)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
por: Lonergan, Liam, et al.
Publicado: (2024)
por: Lonergan, Liam, et al.
Publicado: (2024)
How good is GPT at writing political speeches for the White House?
por: Savoy, Jacques
Publicado: (2024)
por: Savoy, Jacques
Publicado: (2024)
Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
por: Daouadi, Kheir Eddine, et al.
Publicado: (2024)
por: Daouadi, Kheir Eddine, et al.
Publicado: (2024)
VorTEX: Various overlap ratio for Target speech EXtraction
por: Oh, Ro-hoon, et al.
Publicado: (2026)
por: Oh, Ro-hoon, et al.
Publicado: (2026)
Forensic deepfake audio detection using segmental speech features
por: Yang, Tianle, et al.
Publicado: (2025)
por: Yang, Tianle, et al.
Publicado: (2025)
Decoding the decoder: Contextual sequence-to-sequence modeling for intracortical speech decoding
por: Olak, Michal, et al.
Publicado: (2026)
por: Olak, Michal, et al.
Publicado: (2026)
Tag and correct: high precision post-editing approach to correction of speech recognition errors
por: Ziętkiewicz, Tomasz
Publicado: (2024)
por: Ziętkiewicz, Tomasz
Publicado: (2024)
Tracking the emergence of linguistic structure in self-supervised models learning from speech
por: Kloots, Marianne de Heer, et al.
Publicado: (2026)
por: Kloots, Marianne de Heer, et al.
Publicado: (2026)
A Code Comprehension Benchmark for Large Language Models for Code
por: Havare, Jayant, et al.
Publicado: (2025)
por: Havare, Jayant, et al.
Publicado: (2025)
BBPE16: UTF-16-based byte-level byte-pair encoding for improved multilingual speech recognition
por: Kim, Hyunsik, et al.
Publicado: (2026)
por: Kim, Hyunsik, et al.
Publicado: (2026)
Classification is a RAG problem: A case study on hate speech detection
por: Willats, Richard, et al.
Publicado: (2025)
por: Willats, Richard, et al.
Publicado: (2025)
Development and multi-center evaluation of domain-adapted speech recognition for human-AI teaming in real-world gastrointestinal endoscopy
por: Yang, Ruijie, et al.
Publicado: (2026)
por: Yang, Ruijie, et al.
Publicado: (2026)
From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning
por: Neema, Nishit, et al.
Publicado: (2025)
por: Neema, Nishit, et al.
Publicado: (2025)
Can we trust AI to detect healthy multilingual English speakers among the cognitively impaired cohort in the UK? An investigation using real-world conversational speech
por: Pahar, Madhurananda, et al.
Publicado: (2026)
por: Pahar, Madhurananda, et al.
Publicado: (2026)
Enhancing nonnative speech perception and production through an AI-powered application
por: Georgiou, Georgios P.
Publicado: (2025)
por: Georgiou, Georgios P.
Publicado: (2025)
Moshi: a speech-text foundation model for real-time dialogue
por: Défossez, Alexandre, et al.
Publicado: (2024)
por: Défossez, Alexandre, et al.
Publicado: (2024)
What does it take to get state of the art in simultaneous speech-to-speech translation?
por: Wilmet, Vincent, et al.
Publicado: (2024)
por: Wilmet, Vincent, et al.
Publicado: (2024)
Investigating the impact of 2D gesture representation on co-speech gesture generation
por: Guichoux, Teo, et al.
Publicado: (2024)
por: Guichoux, Teo, et al.
Publicado: (2024)
NeuroVoz: a Castillian Spanish corpus of parkinsonian speech
por: Mendes-Laureano, Janaína, et al.
Publicado: (2024)
por: Mendes-Laureano, Janaína, et al.
Publicado: (2024)
Assessing and Mitigating Data Memorization Risks in Fine-Tuned Large Language Models
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
Do self-supervised speech and language models extract similar representations as human brain?
por: Chen, Peili, et al.
Publicado: (2023)
por: Chen, Peili, et al.
Publicado: (2023)
Exploring the traditional NMT model and Large Language Model for chat translation
por: Yang, Jinlong, et al.
Publicado: (2024)
por: Yang, Jinlong, et al.
Publicado: (2024)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
por: C, Anandh, et al.
Publicado: (2025)
por: C, Anandh, et al.
Publicado: (2025)
InstructAudio: Unified speech and music generation with natural language instruction
por: Qiang, Chunyu, et al.
Publicado: (2025)
por: Qiang, Chunyu, et al.
Publicado: (2025)
A unified front-end framework for English text-to-speech synthesis
por: Ying, Zelin, et al.
Publicado: (2023)
por: Ying, Zelin, et al.
Publicado: (2023)
Persuasion Games using Large Language Models
por: Ramani, Ganesh Prasath, et al.
Publicado: (2024)
por: Ramani, Ganesh Prasath, et al.
Publicado: (2024)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
por: Ogun, Sewade, et al.
Publicado: (2024)
por: Ogun, Sewade, et al.
Publicado: (2024)
Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm
por: Shashidhar, Sarvesh, et al.
Publicado: (2025)
por: Shashidhar, Sarvesh, et al.
Publicado: (2025)
Ejemplares similares
-
Dynamic Multi-Expert Projectors with Stabilized Routing for Multilingual Speech Recognition
por: Pandey, Isha, et al.
Publicado: (2026) -
A2TTS: TTS for Low Resource Indian Languages
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025) -
Acoustic and perceptual differences between standard and accented speech and their voice clones
por: Yang, Tianle, et al.
Publicado: (2026) -
Addressing speaker gender bias in large scale speech translation systems
por: Bansal, Shubham, et al.
Publicado: (2025) -
Part-of-speech tagging for Nagamese Language using CRF
por: Shohe, Alovi N, et al.
Publicado: (2025)