Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Yichao, Zhang, Zhirui, Yue, Linan, Huang, Xu, Zhang, Yuqing, Xu, Tong, Xu, Linli, Chen, Enhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
von: Li, Yingting, et al.
Veröffentlicht: (2024)
von: Li, Yingting, et al.
Veröffentlicht: (2024)
A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
von: Xu, Jing, et al.
Veröffentlicht: (2024)
von: Xu, Jing, et al.
Veröffentlicht: (2024)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
Task-Agnostic Structured Pruning of Speech Representation Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2023)
von: Wang, Haoyu, et al.
Veröffentlicht: (2023)
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
BATON: Aligning Text-to-Audio Model with Human Preference Feedback
von: Liao, Huan, et al.
Veröffentlicht: (2024)
von: Liao, Huan, et al.
Veröffentlicht: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
Unified Pathological Speech Analysis with Prompt Tuning
von: Yang, Fei, et al.
Veröffentlicht: (2024)
von: Yang, Fei, et al.
Veröffentlicht: (2024)
Self-Powered LLM Modality Expansion for Large Speech-Text Models
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
von: Yu, Tengfei, et al.
Veröffentlicht: (2024)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Braille-to-Speech Generator: Audio Generation Based on Joint Fine-Tuning of CLIP and Fastspeech2
von: Xu, Chun, et al.
Veröffentlicht: (2024)
von: Xu, Chun, et al.
Veröffentlicht: (2024)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement
von: Pandey, Ashutosh, et al.
Veröffentlicht: (2024)
von: Pandey, Ashutosh, et al.
Veröffentlicht: (2024)
HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech
von: Dong, Zhongren, et al.
Veröffentlicht: (2024)
von: Dong, Zhongren, et al.
Veröffentlicht: (2024)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition
von: Chen, Weidong, et al.
Veröffentlicht: (2023)
von: Chen, Weidong, et al.
Veröffentlicht: (2023)
Bayesian Speech Synthesizers Can Learn from Multiple Teachers
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
von: Han, Bing, et al.
Veröffentlicht: (2024)
von: Han, Bing, et al.
Veröffentlicht: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024) -
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
von: Xie, Jingran, et al.
Veröffentlicht: (2025) -
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
von: Li, Zhu, et al.
Veröffentlicht: (2025) -
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
von: Zeng, Aohan, et al.
Veröffentlicht: (2024) -
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
von: Li, Yingting, et al.
Veröffentlicht: (2024)