FastTalker: Jointly Generating Speech and Conversational Gestures from Text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Zixin, Zhang, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
Conformer-based Ultrasound-to-Speech Conversion
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
Intelligent Text-Conditioned Music Generation
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit
von: Wang, Yutian, et al.
Veröffentlicht: (2024)
von: Wang, Yutian, et al.
Veröffentlicht: (2024)
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
von: Cui, Yang, et al.
Veröffentlicht: (2025)
von: Cui, Yang, et al.
Veröffentlicht: (2025)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
von: Kim, Haven, et al.
Veröffentlicht: (2025)
von: Kim, Haven, et al.
Veröffentlicht: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
von: Wang, Anna, et al.
Veröffentlicht: (2024)
von: Wang, Anna, et al.
Veröffentlicht: (2024)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
von: Lei, Ke, et al.
Veröffentlicht: (2026)
von: Lei, Ke, et al.
Veröffentlicht: (2026)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
Listening Between the Lines: Synthetic Speech Detection Disregarding Verbal Content
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
von: Li, Jizhen, et al.
Veröffentlicht: (2024)
von: Li, Jizhen, et al.
Veröffentlicht: (2024)
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
von: Gong, Jingyao
Veröffentlicht: (2026)
von: Gong, Jingyao
Veröffentlicht: (2026)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
von: Lin, Jinwei
Veröffentlicht: (2024)
von: Lin, Jinwei
Veröffentlicht: (2024)
Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
Dance-to-Music Generation with Encoder-based Textual Inversion
von: Li, Sifei, et al.
Veröffentlicht: (2024)
von: Li, Sifei, et al.
Veröffentlicht: (2024)
StereoFoley: Object-Aware Stereo Audio Generation from Video
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
FGAS: Fixed Decoder Network-Based Audio Steganography with Adversarial Perturbation Generation
von: Yan, Jialin, et al.
Veröffentlicht: (2025)
von: Yan, Jialin, et al.
Veröffentlicht: (2025)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
von: Ye, Zhen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
von: Li, Fengjin, et al.
Veröffentlicht: (2025) -
Conformer-based Ultrasound-to-Speech Conversion
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025) -
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024) -
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025) -
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)