SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Biner, Burak Can, Sofian, Farrin Marouf, Karakaş, Umur Berkay, Ceylan, Duygu, Erdem, Erkut, Erdem, Aykut |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026)
by: Kizil, Muhammed Burak, et al.
Published: (2026)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)
by: Kizil, Muhammed Burak, et al.
Published: (2025)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024)
by: Anees, Abdul Basit, et al.
Published: (2024)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023)
by: Ali, Moayed Haji, et al.
Published: (2023)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)
by: Ercan, Burak, et al.
Published: (2024)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
by: Ekin, Yigit, et al.
Published: (2024)
by: Ekin, Yigit, et al.
Published: (2024)
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025)
by: Çapuk, Hakan, et al.
Published: (2025)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)
by: Karanfil, Enes, et al.
Published: (2025)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
by: Sanli, Enes, et al.
Published: (2025)
by: Sanli, Enes, et al.
Published: (2025)
Variational Control for Guidance in Diffusion Models
by: Pandey, Kushagra, et al.
Published: (2025)
by: Pandey, Kushagra, et al.
Published: (2025)
Control-Augmented Autoregressive Diffusion for Data Assimilation
by: Srivastava, Prakhar, et al.
Published: (2025)
by: Srivastava, Prakhar, et al.
Published: (2025)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026)
by: Bond, Andrew, et al.
Published: (2026)
Hierarchical Variational Policies for Reward-Guided Diffusion
by: Pandey, Kushagra, et al.
Published: (2026)
by: Pandey, Kushagra, et al.
Published: (2026)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025)
by: Bond, Andrew, et al.
Published: (2025)
Edit2Restore:Few-Shot Image Restoration via Parameter-Efficient Adaptation of Pre-trained Editing Models
by: Yılmaz, M. Akın, et al.
Published: (2026)
by: Yılmaz, M. Akın, et al.
Published: (2026)
MD-ProjTex: Texturing 3D Shapes with Multi-Diffusion Projection
by: Yildirim, Ahmet Burak, et al.
Published: (2025)
by: Yildirim, Ahmet Burak, et al.
Published: (2025)
Infrared Domain Adaptation with Zero-Shot Quantization
by: Sevsay, Burak, et al.
Published: (2024)
by: Sevsay, Burak, et al.
Published: (2024)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
Leveraging Image Editing Foundation Models for Data-Efficient CT Metal Artifact Reduction
by: Emirdagi, Ahmet Rasim, et al.
Published: (2026)
by: Emirdagi, Ahmet Rasim, et al.
Published: (2026)
GeoFusionLRM: Geometry-Aware Self-Correction for Consistent 3D Reconstruction
by: Yildirim, Ahmet Burak, et al.
Published: (2026)
by: Yildirim, Ahmet Burak, et al.
Published: (2026)
FuseFormer: A Transformer for Visual and Thermal Image Fusion
by: Erdogan, Aytekin, et al.
Published: (2024)
by: Erdogan, Aytekin, et al.
Published: (2024)
ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
by: Kara, Ozgur, et al.
Published: (2025)
by: Kara, Ozgur, et al.
Published: (2025)
Edit2Interp: Adapting Image Foundation Models from Spatial Editing to Video Frame Interpolation with Few-Shot Learning
by: Rahimi, Nasrin, et al.
Published: (2026)
by: Rahimi, Nasrin, et al.
Published: (2026)
Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models
by: Jiang, Rui, et al.
Published: (2025)
by: Jiang, Rui, et al.
Published: (2025)
FewMMBench: A Benchmark for Multimodal Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2026)
by: Dogan, Mustafa, et al.
Published: (2026)
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
by: Lee, Dohun, et al.
Published: (2026)
by: Lee, Dohun, et al.
Published: (2026)
VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors
by: Koo, Juil, et al.
Published: (2025)
by: Koo, Juil, et al.
Published: (2025)
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
by: Attaiki, Souhaib, et al.
Published: (2024)
by: Attaiki, Souhaib, et al.
Published: (2024)
Object and Relation Centric Representations for Push Effect Prediction
by: Tekden, Ahmet E., et al.
Published: (2021)
by: Tekden, Ahmet E., et al.
Published: (2021)
Sequential Compositional Generalization in Multimodal Models
by: Yagcioglu, Semih, et al.
Published: (2024)
by: Yagcioglu, Semih, et al.
Published: (2024)
EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model
by: Kim, Kunho, et al.
Published: (2026)
by: Kim, Kunho, et al.
Published: (2026)
Boosting Camera Motion Control for Video Diffusion Transformers
by: Cheong, Soon Yau, et al.
Published: (2024)
by: Cheong, Soon Yau, et al.
Published: (2024)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
by: Jeong, Hyeonho, et al.
Published: (2024)
by: Jeong, Hyeonho, et al.
Published: (2024)
AdaptiveDrag: Semantic-Driven Dragging on Diffusion-Based Image Editing
by: Chen, DuoSheng, et al.
Published: (2024)
by: Chen, DuoSheng, et al.
Published: (2024)
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
by: Lee, Dohun, et al.
Published: (2025)
by: Lee, Dohun, et al.
Published: (2025)
How to Augment for Atmospheric Turbulence Effects on Thermal Adapted Object Detection Models?
by: Uzun, Engin, et al.
Published: (2024)
by: Uzun, Engin, et al.
Published: (2024)
Similar Items
-
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026) -
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025) -
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024) -
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023) -
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)