StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Jingyue, Yang, Qihui, Chen, Fei Yueh, McAuley, Julian, Leistikow, Randal, Cook, Perry R., Zang, Yongyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search
von: Yang, Qihui, et al.
Veröffentlicht: (2025)
von: Yang, Qihui, et al.
Veröffentlicht: (2025)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
von: Paek, Nathan, et al.
Veröffentlicht: (2025)
von: Paek, Nathan, et al.
Veröffentlicht: (2025)
Efficient Vocal Source Separation Through Windowed Sink Attention
von: Benetatos, Christodoulos, et al.
Veröffentlicht: (2025)
von: Benetatos, Christodoulos, et al.
Veröffentlicht: (2025)
Composer Vector: Style-steering Symbolic Music Generation in a Latent Space
von: Jiang, Xunyi, et al.
Veröffentlicht: (2026)
von: Jiang, Xunyi, et al.
Veröffentlicht: (2026)
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
Are you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
von: Yang, Qihui, et al.
Veröffentlicht: (2025)
von: Yang, Qihui, et al.
Veröffentlicht: (2025)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
Smule Renaissance Small: Efficient General-Purpose Vocal Restoration
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
SingFake: Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
von: Zang, Yongyi, et al.
Veröffentlicht: (2023)
StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu
Veröffentlicht: (2025)
von: Kim, Nam-Gyu
Veröffentlicht: (2025)
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck
von: Hu, Zhetao, et al.
Veröffentlicht: (2026)
von: Hu, Zhetao, et al.
Veröffentlicht: (2026)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation
von: Guo, Wenxiang, et al.
Veröffentlicht: (2025)
von: Guo, Wenxiang, et al.
Veröffentlicht: (2025)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
von: Huang, Jingyue, et al.
Veröffentlicht: (2025)
von: Huang, Jingyue, et al.
Veröffentlicht: (2025)
Serenade: A Singing Style Conversion Framework Based On Audio Infilling
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
von: Violeta, Lester Phillip, et al.
Veröffentlicht: (2025)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
von: Zhang, You, et al.
Veröffentlicht: (2024)
von: Zhang, You, et al.
Veröffentlicht: (2024)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
GSound-SIR: A Spatial Impulse Response Ray-Tracing and High-order Ambisonic Auralization Python Toolkit
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
Structure-Aware Piano Accompaniment via Style Planning and Dataset-Aligned Pattern Retrieval
von: Zang, Wanyu, et al.
Veröffentlicht: (2026)
von: Zang, Wanyu, et al.
Veröffentlicht: (2026)
Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
von: Kang, Wonjune, et al.
Veröffentlicht: (2025)
von: Kang, Wonjune, et al.
Veröffentlicht: (2025)
Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
von: Huang, Qiaochu, et al.
Veröffentlicht: (2024)
von: Huang, Qiaochu, et al.
Veröffentlicht: (2024)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
von: Kim, Haven, et al.
Veröffentlicht: (2025)
von: Kim, Haven, et al.
Veröffentlicht: (2025)
MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset
von: Xu, Weihan, et al.
Veröffentlicht: (2024)
von: Xu, Weihan, et al.
Veröffentlicht: (2024)
S2Cap: A Benchmark and a Baseline for Singing Style Captioning
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2024)
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
von: Long, Phillip, et al.
Veröffentlicht: (2024)
von: Long, Phillip, et al.
Veröffentlicht: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
MuseCPBench: an Empirical Study of Music Editing Methods through Music Context Preservation
von: Vishe, Yash, et al.
Veröffentlicht: (2025)
von: Vishe, Yash, et al.
Veröffentlicht: (2025)
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search
von: Yang, Qihui, et al.
Veröffentlicht: (2025) -
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
von: Paek, Nathan, et al.
Veröffentlicht: (2025) -
Efficient Vocal Source Separation Through Windowed Sink Attention
von: Benetatos, Christodoulos, et al.
Veröffentlicht: (2025) -
Composer Vector: Style-steering Symbolic Music Generation in a Latent Space
von: Jiang, Xunyi, et al.
Veröffentlicht: (2026) -
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)