PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Huadai, Luo, Kaicheng, Wang, Wen, Chen, Qian, Sun, Peiwen, Huang, Rongjie, Li, Xiangang, Ye, Jieping, Xue, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
by: Cao, Yuqin, et al.
Published: (2024)
by: Cao, Yuqin, et al.
Published: (2024)
Efficient Face Detection with Audio-Based Region Proposals for Human-Robot Interactions
by: Aris, William, et al.
Published: (2023)
by: Aris, William, et al.
Published: (2023)
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
by: Mingote, Victoria, et al.
Published: (2024)
by: Mingote, Victoria, et al.
Published: (2024)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
by: Liu, Huadai, et al.
Published: (2024)
by: Liu, Huadai, et al.
Published: (2024)
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
by: Yue, Xianghu, et al.
Published: (2024)
by: Yue, Xianghu, et al.
Published: (2024)
OmniAudio: Generating Spatial Audio from 360-Degree Video
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
by: Liu, Huadai, et al.
Published: (2024)
by: Liu, Huadai, et al.
Published: (2024)
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
KunquDB: An Attempt for Speaker Verification in the Chinese Opera Scenario
by: Zhou, Huali, et al.
Published: (2024)
by: Zhou, Huali, et al.
Published: (2024)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
by: Park, Jay, et al.
Published: (2025)
by: Park, Jay, et al.
Published: (2025)
Improvement Of Audiovisual Quality Estimation Using A Nonlinear Autoregressive Exogenous Neural Network And Bitstream Parameters
by: Kossi, Koffi, et al.
Published: (2024)
by: Kossi, Koffi, et al.
Published: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
by: Chen, Xuanjun, et al.
Published: (2025)
by: Chen, Xuanjun, et al.
Published: (2025)
Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2021)
by: Praveen, R. Gnana, et al.
Published: (2021)
Out-Of-Distribution Detection for Audio-visual Generalized Zero-Shot Learning: A General Framework
by: Wen, Liuyuan
Published: (2024)
by: Wen, Liuyuan
Published: (2024)
Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index
by: Ragano, Alessandro, et al.
Published: (2024)
by: Ragano, Alessandro, et al.
Published: (2024)
Audio-Visual Approach For Multimodal Concurrent Speaker Detection
by: Eliav, Amit, et al.
Published: (2024)
by: Eliav, Amit, et al.
Published: (2024)
Cross-modal Cognitive Consensus guided Audio-Visual Segmentation
by: Shi, Zhaofeng, et al.
Published: (2023)
by: Shi, Zhaofeng, et al.
Published: (2023)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
by: Yang, Dongchao, et al.
Published: (2023)
by: Yang, Dongchao, et al.
Published: (2023)
ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
by: Berghi, Davide, et al.
Published: (2026)
by: Berghi, Davide, et al.
Published: (2026)
Multimodal Marvels of Deep Learning in Medical Diagnosis: A Comprehensive Review of COVID-19 Detection
by: Islam, Md Shofiqul, et al.
Published: (2025)
by: Islam, Md Shofiqul, et al.
Published: (2025)
End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments
by: Lin, Meng-Ping, et al.
Published: (2025)
by: Lin, Meng-Ping, et al.
Published: (2025)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
by: Yang, Dongchao, et al.
Published: (2024)
by: Yang, Dongchao, et al.
Published: (2024)
Machine Perceptual Quality: Evaluating the Impact of Severe Lossy Compression on Audio and Image Models
by: Jacobellis, Dan, et al.
Published: (2024)
by: Jacobellis, Dan, et al.
Published: (2024)
SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
by: Tanigawa, Risako, et al.
Published: (2024)
by: Tanigawa, Risako, et al.
Published: (2024)
A multi-modal approach for identifying schizophrenia using cross-modal attention
by: Premananth, Gowtham, et al.
Published: (2023)
by: Premananth, Gowtham, et al.
Published: (2023)
Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
by: Sun, Peiwen, et al.
Published: (2024)
by: Sun, Peiwen, et al.
Published: (2024)
Unified Audio Event Detection
by: Jiang, Yidi, et al.
Published: (2024)
by: Jiang, Yidi, et al.
Published: (2024)
Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
by: Salaj, Ina, et al.
Published: (2025)
by: Salaj, Ina, et al.
Published: (2025)
Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs
by: Jia, Yuhang, et al.
Published: (2025)
by: Jia, Yuhang, et al.
Published: (2025)
MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
by: Liu, Huadai, et al.
Published: (2024)
by: Liu, Huadai, et al.
Published: (2024)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
by: Nguyen, Hong, et al.
Published: (2024)
by: Nguyen, Hong, et al.
Published: (2024)
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
by: Zhong, Zhi, et al.
Published: (2025)
by: Zhong, Zhi, et al.
Published: (2025)
Livestock feeding behaviour: A review on automated systems for ruminant monitoring
by: Chelotti, José, et al.
Published: (2023)
by: Chelotti, José, et al.
Published: (2023)
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
by: Ishikawa, Yuchi, et al.
Published: (2025)
by: Ishikawa, Yuchi, et al.
Published: (2025)
AudioCIL: A Python Toolbox for Audio Class-Incremental Learning with Multiple Scenes
by: Xu, Qisheng, et al.
Published: (2024)
by: Xu, Qisheng, et al.
Published: (2024)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
by: Sulun, Serkan, et al.
Published: (2025)
by: Sulun, Serkan, et al.
Published: (2025)
QuarkAudio Technical Report
by: Liu, Chengwei, et al.
Published: (2025)
by: Liu, Chengwei, et al.
Published: (2025)
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
Similar Items
-
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
by: Cao, Yuqin, et al.
Published: (2024) -
Efficient Face Detection with Audio-Based Region Proposals for Human-Robot Interactions
by: Aris, William, et al.
Published: (2023) -
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
by: Mingote, Victoria, et al.
Published: (2024) -
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
by: Liu, Huadai, et al.
Published: (2024) -
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
by: Liu, Huadai, et al.
Published: (2025)