Images that Sound: Composing Images and Sounds on a Single Canvas
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Ziyang, Geng, Daniel, Owens, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
Sounding that Object: Interactive Object-Aware Image to Audio Generation
by: Li, Tingle, et al.
Published: (2025)
by: Li, Tingle, et al.
Published: (2025)
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
by: Lee, Junwon, et al.
Published: (2024)
by: Lee, Junwon, et al.
Published: (2024)
Continual Audio-Visual Sound Separation
by: Pian, Weiguo, et al.
Published: (2024)
by: Pian, Weiguo, et al.
Published: (2024)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer
by: Wang, Yaoting, et al.
Published: (2023)
by: Wang, Yaoting, et al.
Published: (2023)
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
by: Gramaccioni, Riccardo Fosco, et al.
Published: (2024)
by: Gramaccioni, Riccardo Fosco, et al.
Published: (2024)
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning
by: Singh, Nikhil, et al.
Published: (2023)
by: Singh, Nikhil, et al.
Published: (2023)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
by: Chen, Zihao, et al.
Published: (2024)
by: Chen, Zihao, et al.
Published: (2024)
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
by: Senocak, Arda, et al.
Published: (2024)
by: Senocak, Arda, et al.
Published: (2024)
The Sound of Water: Inferring Physical Properties from Pouring Liquids
by: Bagad, Piyush, et al.
Published: (2024)
by: Bagad, Piyush, et al.
Published: (2024)
Self-Supervised Audio-Visual Soundscape Stylization
by: Li, Tingle, et al.
Published: (2024)
by: Li, Tingle, et al.
Published: (2024)
Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
by: Krishnamurthy, Sudha
Published: (2024)
by: Krishnamurthy, Sudha
Published: (2024)
Read, Watch and Scream! Sound Generation from Text and Video
by: Jeong, Yujin, et al.
Published: (2024)
by: Jeong, Yujin, et al.
Published: (2024)
SemiPL: A Semi-supervised Method for Event Sound Source Localization
by: Li, Yue, et al.
Published: (2024)
by: Li, Yue, et al.
Published: (2024)
Gotta Hear Them All: Towards Sound Source Aware Audio Generation
by: Guo, Wei, et al.
Published: (2024)
by: Guo, Wei, et al.
Published: (2024)
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
by: Kim, Dongjin, et al.
Published: (2024)
by: Kim, Dongjin, et al.
Published: (2024)
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
by: Sung-Bin, Kim, et al.
Published: (2024)
by: Sung-Bin, Kim, et al.
Published: (2024)
FilmComposer: LLM-Driven Music Production for Silent Film Clips
by: Xie, Zhifeng, et al.
Published: (2025)
by: Xie, Zhifeng, et al.
Published: (2025)
Synchformer: Efficient Synchronization from Sparse Cues
by: Iashin, Vladimir, et al.
Published: (2024)
by: Iashin, Vladimir, et al.
Published: (2024)
Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization
by: Jia, Yanhao, et al.
Published: (2025)
by: Jia, Yanhao, et al.
Published: (2025)
A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition
by: Dai, Yusheng, et al.
Published: (2024)
by: Dai, Yusheng, et al.
Published: (2024)
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
by: Fang, Pengjun, et al.
Published: (2026)
by: Fang, Pengjun, et al.
Published: (2026)
AudioX: A Unified Framework for Anything-to-Audio Generation
by: Tian, Zeyue, et al.
Published: (2025)
by: Tian, Zeyue, et al.
Published: (2025)
Audio-Visual Instance Segmentation
by: Guo, Ruohao, et al.
Published: (2023)
by: Guo, Ruohao, et al.
Published: (2023)
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
by: Hayakawa, Akio, et al.
Published: (2024)
by: Hayakawa, Akio, et al.
Published: (2024)
Improving Multimodal Learning with Multi-Loss Gradient Modulation
by: Kontras, Konstantinos, et al.
Published: (2024)
by: Kontras, Konstantinos, et al.
Published: (2024)
AMuSE: Adaptive Multimodal Analysis for Speaker Emotion Recognition in Group Conversations
by: Devulapally, Naresh Kumar, et al.
Published: (2024)
by: Devulapally, Naresh Kumar, et al.
Published: (2024)
Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization
by: Guo, Yuxin, et al.
Published: (2024)
by: Guo, Yuxin, et al.
Published: (2024)
Collaborative Hybrid Propagator for Temporal Misalignment in Audio-Visual Segmentation
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation
by: Su, Kun, et al.
Published: (2024)
by: Su, Kun, et al.
Published: (2024)
Audio-visual Generalized Zero-shot Learning the Easy Way
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation
by: Kim, Gwanghyun, et al.
Published: (2024)
by: Kim, Gwanghyun, et al.
Published: (2024)
FairSSD: Understanding Bias in Synthetic Speech Detectors
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
Sequential Contrastive Audio-Visual Learning
by: Tsiamas, Ioannis, et al.
Published: (2024)
by: Tsiamas, Ioannis, et al.
Published: (2024)
Semantic Grouping Network for Audio Source Separation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Similar Items
-
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024) -
Sounding that Object: Interactive Object-Aware Image to Audio Generation
by: Li, Tingle, et al.
Published: (2025) -
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
by: Lee, Junwon, et al.
Published: (2024) -
Continual Audio-Visual Sound Separation
by: Pian, Weiguo, et al.
Published: (2024) -
Multi-scale Multi-instance Visual Sound Localization and Segmentation
by: Mo, Shentong, et al.
Published: (2024)