HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Niu, Xinlei, Zhang, Jing, Martin, Charles Patrick |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
von: Lin, Jinwei
Veröffentlicht: (2024)
von: Lin, Jinwei
Veröffentlicht: (2024)
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
von: Huang, Lian, et al.
Veröffentlicht: (2024)
von: Huang, Lian, et al.
Veröffentlicht: (2024)
Resource-Efficient Reference-Free Evaluation of Audio Captions
von: Mahfuz, Rehana, et al.
Veröffentlicht: (2024)
von: Mahfuz, Rehana, et al.
Veröffentlicht: (2024)
Efficient Video to Audio Mapper with Visual Scene Detection
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
Neural Style Transfer for Audio Spectograms
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
von: Zhang, You, et al.
Veröffentlicht: (2024)
von: Zhang, You, et al.
Veröffentlicht: (2024)
Intelligent Text-Conditioned Music Generation
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
von: Xie, Zhouyao, et al.
Veröffentlicht: (2024)
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
von: Liu, Shengqiang, et al.
Veröffentlicht: (2024)
von: Liu, Shengqiang, et al.
Veröffentlicht: (2024)
MART: Learning Hierarchical Music Audio Representations with Part-Whole Transformer
von: Yao, Dong, et al.
Veröffentlicht: (2023)
von: Yao, Dong, et al.
Veröffentlicht: (2023)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Attentive-based Multi-level Feature Fusion for Voice Disorder Diagnosis
von: Shen, Lipeng, et al.
Veröffentlicht: (2024)
von: Shen, Lipeng, et al.
Veröffentlicht: (2024)
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
Building Audio-Visual Digital Twins with Smartphones
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
FGAS: Fixed Decoder Network-Based Audio Steganography with Adversarial Perturbation Generation
von: Yan, Jialin, et al.
Veröffentlicht: (2025)
von: Yan, Jialin, et al.
Veröffentlicht: (2025)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
von: Lei, Ke, et al.
Veröffentlicht: (2026)
von: Lei, Ke, et al.
Veröffentlicht: (2026)
Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
von: Huang, Qiaochu, et al.
Veröffentlicht: (2024)
von: Huang, Qiaochu, et al.
Veröffentlicht: (2024)
LoVA: Long-form Video-to-Audio Generation
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
von: Shimada, Kazuki, et al.
Veröffentlicht: (2024)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025) -
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
von: Li, Fengjin, et al.
Veröffentlicht: (2025) -
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing
von: Niu, Xinlei, et al.
Veröffentlicht: (2025) -
ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
von: Lin, Jinwei
Veröffentlicht: (2024) -
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)