BATON: Aligning Text-to-Audio Model with Human Preference Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liao, Huan, Han, Haonan, Yang, Kai, Du, Tianjiao, Yang, Rui, Xu, Zunnan, Xu, Qinmei, Liu, Jingquan, Lu, Jiasheng, Li, Xiu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Aligning Audio Captions with Human Preferences
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
Towards Weakly Supervised Text-to-Audio Grounding
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Aligning Text-to-Music Evaluation with Human Preferences
von: Huang, Yichen, et al.
Veröffentlicht: (2025)
von: Huang, Yichen, et al.
Veröffentlicht: (2025)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
von: Liu, Wuyang, et al.
Veröffentlicht: (2023)
von: Liu, Wuyang, et al.
Veröffentlicht: (2023)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
Text2Move: Text-to-moving sound generation via trajectory prediction and temporal alignment
von: Liu, Yunyi, et al.
Veröffentlicht: (2025)
von: Liu, Yunyi, et al.
Veröffentlicht: (2025)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
SpeechAlign: Aligning Speech Generation to Human Preferences
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
von: Yin, Han, et al.
Veröffentlicht: (2024)
von: Yin, Han, et al.
Veröffentlicht: (2024)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
von: Gu, Yi, et al.
Veröffentlicht: (2026)
von: Gu, Yi, et al.
Veröffentlicht: (2026)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
MusicRL: Aligning Music Generation to Human Preferences
von: Cideron, Geoffrey, et al.
Veröffentlicht: (2024)
von: Cideron, Geoffrey, et al.
Veröffentlicht: (2024)
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation
von: Foo, Leonardo Haw-Yang, et al.
Veröffentlicht: (2026)
von: Foo, Leonardo Haw-Yang, et al.
Veröffentlicht: (2026)
A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
Exploring Differences between Human Perception and Model Inference in Audio Event Recognition
von: Tan, Yizhou, et al.
Veröffentlicht: (2024)
von: Tan, Yizhou, et al.
Veröffentlicht: (2024)
Aligning Generative Music AI with Human Preferences: Methods and Challenges
von: Herremans, Dorien, et al.
Veröffentlicht: (2025)
von: Herremans, Dorien, et al.
Veröffentlicht: (2025)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
von: Yang, Mu, et al.
Veröffentlicht: (2024)
von: Yang, Mu, et al.
Veröffentlicht: (2024)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
von: Xin, Detai, et al.
Veröffentlicht: (2026)
von: Xin, Detai, et al.
Veröffentlicht: (2026)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
Cacophony: An Improved Contrastive Audio-Text Model
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
AudioSpa: Spatializing Sound Events with Text
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
Enhancing Crowdsourced Audio for Text-to-Speech Models
von: Giraldo, José, et al.
Veröffentlicht: (2024)
von: Giraldo, José, et al.
Veröffentlicht: (2024)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Aligning Audio Captions with Human Preferences
von: Hegde, Kartik, et al.
Veröffentlicht: (2025) -
Towards Weakly Supervised Text-to-Audio Grounding
von: Xu, Xuenan, et al.
Veröffentlicht: (2024) -
Aligning Text-to-Music Evaluation with Human Preferences
von: Huang, Yichen, et al.
Veröffentlicht: (2025) -
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
von: Liu, Wuyang, et al.
Veröffentlicht: (2023) -
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)