Saved in:
| Main Authors: | Yang, Xiaoda, Zhang, Majun, Pan, Changhao, Huang, Nick, Yuguang, Yang, Zhuo, Fan, Zhou, Pengfei, Zhou, Jin, Shan, Sizhe, Yang, Shan, Yang, Miles, You, Yang, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.01809 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
by: Shan, Sizhe, et al.
Published: (2025)
by: Shan, Sizhe, et al.
Published: (2025)
TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
by: Yang, Ziyue, et al.
Published: (2026)
by: Yang, Ziyue, et al.
Published: (2026)
ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks
by: Xu, Jiayang, et al.
Published: (2026)
by: Xu, Jiayang, et al.
Published: (2026)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
by: Chen, Peikun, et al.
Published: (2024)
by: Chen, Peikun, et al.
Published: (2024)
SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation
by: Wang, Hongrui, et al.
Published: (2026)
by: Wang, Hongrui, et al.
Published: (2026)
DanceChat: Large Language Model-Guided Music-to-Dance Generation
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis
by: Hu, Xintong, et al.
Published: (2025)
by: Hu, Xintong, et al.
Published: (2025)
CoheDancers: Enhancing Interactive Group Dance Generation through Music-Driven Coherence Decomposition
by: Yang, Kaixing, et al.
Published: (2024)
by: Yang, Kaixing, et al.
Published: (2024)
Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty
by: Yang, Zhou, et al.
Published: (2026)
by: Yang, Zhou, et al.
Published: (2026)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
by: Yao, Jixun, et al.
Published: (2024)
by: Yao, Jixun, et al.
Published: (2024)
Exploring Multi-Modal Control in Music-Driven Dance Generation
by: Li, Ronghui, et al.
Published: (2024)
by: Li, Ronghui, et al.
Published: (2024)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
by: Yao, Jixun, et al.
Published: (2025)
by: Yao, Jixun, et al.
Published: (2025)
Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
by: Dai, Congren, et al.
Published: (2025)
by: Dai, Congren, et al.
Published: (2025)
EnchantDance: Unveiling the Potential of Music-Driven Dance Movement
by: Han, Bo, et al.
Published: (2023)
by: Han, Bo, et al.
Published: (2023)
PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders
by: Pan, Yu, et al.
Published: (2024)
by: Pan, Yu, et al.
Published: (2024)
Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation
by: Lyu, Guangtao, et al.
Published: (2025)
by: Lyu, Guangtao, et al.
Published: (2025)
AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
Dance-to-Music Generation with Encoder-based Textual Inversion
by: Li, Sifei, et al.
Published: (2024)
by: Li, Sifei, et al.
Published: (2024)
Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
by: Li, Xiaojie, et al.
Published: (2025)
by: Li, Xiaojie, et al.
Published: (2025)
Acoustic Overspecification in Electronic Dance Music Taxonomy
by: Xu, Weilun, et al.
Published: (2025)
by: Xu, Weilun, et al.
Published: (2025)
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
by: Yang, Qian, et al.
Published: (2024)
by: Yang, Qian, et al.
Published: (2024)
Cross-Modal Denoising: A Novel Training Paradigm for Enhancing Speech-Image Retrieval
by: Zhou, Lifeng, et al.
Published: (2024)
by: Zhou, Lifeng, et al.
Published: (2024)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints
by: Meng, Hao, et al.
Published: (2026)
by: Meng, Hao, et al.
Published: (2026)
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
by: Li, You, et al.
Published: (2026)
by: Li, You, et al.
Published: (2026)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
by: Pan, Yu, et al.
Published: (2024)
by: Pan, Yu, et al.
Published: (2024)
GMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer Learning for Speech Emotion Recognition
by: Pan, Yu, et al.
Published: (2024)
by: Pan, Yu, et al.
Published: (2024)
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
by: Jin, Jiawei, et al.
Published: (2025)
by: Jin, Jiawei, et al.
Published: (2025)
ASAudio: A Survey of Advanced Spatial Audio Research
by: Zhu, Zhiyuan, et al.
Published: (2025)
by: Zhu, Zhiyuan, et al.
Published: (2025)
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
Music2Fail: Transfer Music to Failed Recorder Style
by: Leong, Chon In, et al.
Published: (2024)
by: Leong, Chon In, et al.
Published: (2024)
From Speech to Profile: A Protocol-Driven LLM Agent for Psychological Profile Generation
by: Yang, Xingjian, et al.
Published: (2026)
by: Yang, Xingjian, et al.
Published: (2026)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
by: Chen, Weidong, et al.
Published: (2025)
by: Chen, Weidong, et al.
Published: (2025)
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
by: Yang, Yuguang, et al.
Published: (2024)
by: Yang, Yuguang, et al.
Published: (2024)
Detecting Musical Deepfakes
by: Sunday, Nick
Published: (2025)
by: Sunday, Nick
Published: (2025)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
by: Zuo, Jialong, et al.
Published: (2025)
by: Zuo, Jialong, et al.
Published: (2025)
The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models
by: Li, Jiajia, et al.
Published: (2024)
by: Li, Jiajia, et al.
Published: (2024)
MusER: Musical Element-Based Regularization for Generating Symbolic Music with Emotion
by: Ji, Shulei, et al.
Published: (2023)
by: Ji, Shulei, et al.
Published: (2023)
Similar Items
-
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
by: Shan, Sizhe, et al.
Published: (2025) -
TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
by: Yang, Ziyue, et al.
Published: (2026) -
ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks
by: Xu, Jiayang, et al.
Published: (2026) -
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
by: Chen, Peikun, et al.
Published: (2024) -
SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation
by: Wang, Hongrui, et al.
Published: (2026)