Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Qingkai, Zhang, Shaolei, Ma, Zhengrui, Zhang, Min, Feng, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
von: Ma, Zhengrui, et al.
Veröffentlicht: (2024)
von: Ma, Zhengrui, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Everyday Speech in the Indian Subcontinent
von: P, Utkarsh
Veröffentlicht: (2024)
von: P, Utkarsh
Veröffentlicht: (2024)
Measuring the Accuracy of Automatic Speech Recognition Solutions
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Communication Access Real-Time Translation Through Collaborative Correction of Automatic Speech Recognition
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2025)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
Syllable based DNN-HMM Cantonese Speech to Text System
von: Wong, Timothy, et al.
Veröffentlicht: (2024)
von: Wong, Timothy, et al.
Veröffentlicht: (2024)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
von: Lee, Yoonhyung, et al.
Veröffentlicht: (2026)
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
von: Dua, Karan, et al.
Veröffentlicht: (2025)
von: Dua, Karan, et al.
Veröffentlicht: (2025)
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Quantifying the effect of speech pathology on automatic and human speaker verification
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit
von: Soni, Aniket Abhishek
Veröffentlicht: (2025)
von: Soni, Aniket Abhishek
Veröffentlicht: (2025)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
von: Cheripally, Sowmya
Veröffentlicht: (2024)
von: Cheripally, Sowmya
Veröffentlicht: (2024)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset
von: Marie, Ambre, et al.
Veröffentlicht: (2025)
von: Marie, Ambre, et al.
Veröffentlicht: (2025)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
von: Bhadra, Dipayan, et al.
Veröffentlicht: (2025)
von: Bhadra, Dipayan, et al.
Veröffentlicht: (2025)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
von: Salvi, Giampiero
Veröffentlicht: (2024)
von: Salvi, Giampiero
Veröffentlicht: (2024)
Framework for Curating Speech Datasets and Evaluating ASR Systems: A Case Study for Polish
von: Junczyk, Michał
Veröffentlicht: (2024)
von: Junczyk, Michał
Veröffentlicht: (2024)
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
von: Chen, Zhehuai, et al.
Veröffentlicht: (2024)
von: Chen, Zhehuai, et al.
Veröffentlicht: (2024)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
von: Fang, Qingkai, et al.
Veröffentlicht: (2024) -
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
von: Fang, Qingkai, et al.
Veröffentlicht: (2024) -
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025) -
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024) -
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
von: Ma, Zhengrui, et al.
Veröffentlicht: (2024)