Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Motamed, Saman, Van Gansbeke, Wouter, Van Gool, Luc |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
von: Motamed, Saman, et al.
Veröffentlicht: (2023)
von: Motamed, Saman, et al.
Veröffentlicht: (2023)
A Simple Latent Diffusion Approach for Panoptic Segmentation and Mask Inpainting
von: Van Gansbeke, Wouter, et al.
Veröffentlicht: (2024)
von: Van Gansbeke, Wouter, et al.
Veröffentlicht: (2024)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
von: Motamed, Saman, et al.
Veröffentlicht: (2025)
von: Motamed, Saman, et al.
Veröffentlicht: (2025)
A Simple and Generalist Approach for Panoptic Segmentation
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2024)
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2024)
VOID: Video Object and Interaction Deletion
von: Motamed, Saman, et al.
Veröffentlicht: (2026)
von: Motamed, Saman, et al.
Veröffentlicht: (2026)
A Unified and Interpretable Emotion Representation and Expression Generation
von: Paskaleva, Reni, et al.
Veröffentlicht: (2024)
von: Paskaleva, Reni, et al.
Veröffentlicht: (2024)
Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector
von: Fu, Yuqian, et al.
Veröffentlicht: (2024)
von: Fu, Yuqian, et al.
Veröffentlicht: (2024)
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
Stereo Risk: A Continuous Modeling Approach to Stereo Matching
von: Liu, Ce, et al.
Veröffentlicht: (2024)
von: Liu, Ce, et al.
Veröffentlicht: (2024)
Adversarial Dependence Minimization
von: De Plaen, Pierre-François, et al.
Veröffentlicht: (2025)
von: De Plaen, Pierre-François, et al.
Veröffentlicht: (2025)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
von: Shahbazi, Mohamad, et al.
Veröffentlicht: (2024)
von: Shahbazi, Mohamad, et al.
Veröffentlicht: (2024)
Zero-Shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
von: Cao, Cong, et al.
Veröffentlicht: (2024)
von: Cao, Cong, et al.
Veröffentlicht: (2024)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
von: Broedermannn, Tim, et al.
Veröffentlicht: (2025)
von: Broedermannn, Tim, et al.
Veröffentlicht: (2025)
FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
von: Markov, Mario, et al.
Veröffentlicht: (2025)
von: Markov, Mario, et al.
Veröffentlicht: (2025)
B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation
von: Markov, Mario, et al.
Veröffentlicht: (2026)
von: Markov, Mario, et al.
Veröffentlicht: (2026)
Probabilistic Sampling of Balanced K-Means using Adiabatic Quantum Computing
von: Zaech, Jan-Nico, et al.
Veröffentlicht: (2023)
von: Zaech, Jan-Nico, et al.
Veröffentlicht: (2023)
Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis
von: Mahdi, Mohammad, et al.
Veröffentlicht: (2025)
von: Mahdi, Mohammad, et al.
Veröffentlicht: (2025)
OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs
von: Ailuro, Stefan Maria, et al.
Veröffentlicht: (2026)
von: Ailuro, Stefan Maria, et al.
Veröffentlicht: (2026)
Zero-Shot Medical Phrase Grounding with Off-the-shelf Diffusion Models
von: Vilouras, Konstantinos, et al.
Veröffentlicht: (2024)
von: Vilouras, Konstantinos, et al.
Veröffentlicht: (2024)
Shapley Pruning for Neural Network Compression
von: Adamczewski, Kamil, et al.
Veröffentlicht: (2024)
von: Adamczewski, Kamil, et al.
Veröffentlicht: (2024)
InTraGen: Trajectory-controlled Video Generation for Object Interactions
von: Liu, Zuhao, et al.
Veröffentlicht: (2024)
von: Liu, Zuhao, et al.
Veröffentlicht: (2024)
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
von: Khan, Muhammad Saif Ullah, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Saif Ullah, et al.
Veröffentlicht: (2024)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
Object-Centric Diffusion for Efficient Video Editing
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Efficient Zero-Shot Inpainting with Decoupled Diffusion Guidance
von: Moufad, Badr, et al.
Veröffentlicht: (2025)
von: Moufad, Badr, et al.
Veröffentlicht: (2025)
Unlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2024)
von: Jiang, Eric Hanchen, et al.
Veröffentlicht: (2024)
Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
von: Xu, Chenhui, et al.
Veröffentlicht: (2025)
Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text Guidance
von: Nam, Giung, et al.
Veröffentlicht: (2024)
von: Nam, Giung, et al.
Veröffentlicht: (2024)
Unlearning Concepts from Text-to-Video Diffusion Models
von: Liu, Shiqi, et al.
Veröffentlicht: (2024)
von: Liu, Shiqi, et al.
Veröffentlicht: (2024)
Do generative video models understand physical principles?
von: Motamed, Saman, et al.
Veröffentlicht: (2025)
von: Motamed, Saman, et al.
Veröffentlicht: (2025)
Attention Based Simple Primitives for Open World Compositional Zero-Shot Learning
von: Munir, Ans, et al.
Veröffentlicht: (2024)
von: Munir, Ans, et al.
Veröffentlicht: (2024)
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Masked Extended Attention for Zero-Shot Virtual Try-On In The Wild
von: Orzech, Nadav, et al.
Veröffentlicht: (2024)
von: Orzech, Nadav, et al.
Veröffentlicht: (2024)
Editing Massive Concepts in Text-to-Image Diffusion Models
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
von: Park, Jonggwon, et al.
Veröffentlicht: (2025)
ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation
von: Min, Yunhong, et al.
Veröffentlicht: (2025)
von: Min, Yunhong, et al.
Veröffentlicht: (2025)
Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion
von: Viola, Massimiliano, et al.
Veröffentlicht: (2024)
von: Viola, Massimiliano, et al.
Veröffentlicht: (2024)
Slicedit: Zero-Shot Video Editing With Text-to-Image Diffusion Models Using Spatio-Temporal Slices
von: Cohen, Nathaniel, et al.
Veröffentlicht: (2024)
von: Cohen, Nathaniel, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models
von: Motamed, Saman, et al.
Veröffentlicht: (2023) -
A Simple Latent Diffusion Approach for Panoptic Segmentation and Mask Inpainting
von: Van Gansbeke, Wouter, et al.
Veröffentlicht: (2024) -
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
von: Motamed, Saman, et al.
Veröffentlicht: (2025) -
A Simple and Generalist Approach for Panoptic Segmentation
von: Prisadnikov, Nedyalko, et al.
Veröffentlicht: (2024) -
VOID: Video Object and Interaction Deletion
von: Motamed, Saman, et al.
Veröffentlicht: (2026)