Skip to content
VuFind
  • Login
    • English
    • Deutsch
    • Español
    • Français
    • Italiano
Advanced
  • Cite this
  • Text this
  • Email this
  • Print
  • Export Record
    • Export to RefWorks
    • Export to EndNoteWeb
    • Export to EndNote
  • Save to List
  • Permanent link
Cover Image

Saved in:
Bibliographic Details
Main Authors: Yang, Jianxuan, Guo, Xinyue, Cheng, Zhi, Wang, Kai, Zhang, Lipan, Hu, Jinjie, Ji, Qiang, Cao, Yihua, Meng, Yihao, Cui, Zhaoyue, Liu, Mengmei, Meng, Meng, Luan, Jian
Format: Preprint
Published: 2026
Subjects:
Multimedia
Computer Vision and Pattern Recognition
Sound
Online Access:https://arxiv.org/abs/2604.15086
Tags: Add Tag
No Tags, Be the first to tag this record!
  • Holdings
  • Description
  • Table of Contents
  • Comments
  • Similar Items
  • Staff View

Internet

https://arxiv.org/abs/2604.15086

Similar Items

  • AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
    by: Guo, Xinyue, et al.
    Published: (2025)
  • MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
    by: Yang, Jianxuan, et al.
    Published: (2025)
  • DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
    by: Li, Fu, et al.
    Published: (2025)
  • Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
    by: Huang, Zhiqi, et al.
    Published: (2024)
  • StereoFoley: Object-Aware Stereo Audio Generation from Video
    by: Karchkhadze, Tornike, et al.
    Published: (2025)

Search Options

  • Search History
  • Advanced Search

Find More

  • Browse the Catalog
  • Browse Alphabetically
  • Explore Channels
  • Course Reserves
  • New Items

Need Help?

  • Search Tips
  • Ask a Librarian
  • FAQs