VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Couairon, Paul, Rambour, Clément, Haugeard, Jean-Emmanuel, Thome, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
by: Couairon, Paul, et al.
Published: (2024)
by: Couairon, Paul, et al.
Published: (2024)
JAFAR: Jack up Any Feature at Any Resolution
by: Couairon, Paul, et al.
Published: (2025)
by: Couairon, Paul, et al.
Published: (2025)
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
by: Chambon, Loick, et al.
Published: (2025)
by: Chambon, Loick, et al.
Published: (2025)
Energy Correction Model in the Feature Space for Out-of-Distribution Detection
by: Lafon, Marc, et al.
Published: (2024)
by: Lafon, Marc, et al.
Published: (2024)
GalLoP: Learning Global and Local Prompts for Vision-Language Models
by: Lafon, Marc, et al.
Published: (2024)
by: Lafon, Marc, et al.
Published: (2024)
ViLU: Learning Vision-Language Uncertainties for Failure Prediction
by: Lafon, Marc, et al.
Published: (2025)
by: Lafon, Marc, et al.
Published: (2025)
CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation
by: Lafon, Marc, et al.
Published: (2025)
by: Lafon, Marc, et al.
Published: (2025)
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
by: Seo, Junyoung, et al.
Published: (2025)
by: Seo, Junyoung, et al.
Published: (2025)
DAFTED: Decoupled Asymmetric Fusion of Tabular and Echocardiographic Data for Cardiac Hypertension Diagnosis
by: Stym-Popper, Jérémie, et al.
Published: (2025)
by: Stym-Popper, Jérémie, et al.
Published: (2025)
Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
FREE-Edit: Using Editing-aware Injection in Rectified Flow Models for Zero-shot Image-Driven Video Editing
by: Li, Maomao, et al.
Published: (2026)
by: Li, Maomao, et al.
Published: (2026)
FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing
by: Zhang, Youyuan, et al.
Published: (2024)
by: Zhang, Youyuan, et al.
Published: (2024)
PostEdit: Posterior Sampling for Efficient Zero-Shot Image Editing
by: Tian, Feng, et al.
Published: (2024)
by: Tian, Feng, et al.
Published: (2024)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
by: Hoe, Jiun Tian, et al.
Published: (2025)
by: Hoe, Jiun Tian, et al.
Published: (2025)
S$^2$Edit: Text-Guided Image Editing with Precise Semantic and Spatial Control
by: Liu, Xudong, et al.
Published: (2025)
by: Liu, Xudong, et al.
Published: (2025)
Edit2Interp: Adapting Image Foundation Models from Spatial Editing to Video Frame Interpolation with Few-Shot Learning
by: Rahimi, Nasrin, et al.
Published: (2026)
by: Rahimi, Nasrin, et al.
Published: (2026)
A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing
by: Li, Maomao, et al.
Published: (2023)
by: Li, Maomao, et al.
Published: (2023)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
by: Yang, Zhoufaran, et al.
Published: (2025)
by: Yang, Zhoufaran, et al.
Published: (2025)
OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical Videos
by: Jangir, Ritul, et al.
Published: (2026)
by: Jangir, Ritul, et al.
Published: (2026)
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
by: Somayazulu, Arjun, et al.
Published: (2026)
by: Somayazulu, Arjun, et al.
Published: (2026)
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
by: Xiao, Yicheng, et al.
Published: (2026)
by: Xiao, Yicheng, et al.
Published: (2026)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
Slicedit: Zero-Shot Video Editing With Text-to-Image Diffusion Models Using Spatio-Temporal Slices
by: Cohen, Nathaniel, et al.
Published: (2024)
by: Cohen, Nathaniel, et al.
Published: (2024)
AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
by: Zheng, Haojie, et al.
Published: (2025)
by: Zheng, Haojie, et al.
Published: (2025)
ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks
by: Xu, Jiayang, et al.
Published: (2026)
by: Xu, Jiayang, et al.
Published: (2026)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
by: Zhong, Yangyang, et al.
Published: (2025)
by: Zhong, Yangyang, et al.
Published: (2025)
Edit Fidelity Field: Semantics-Aware Region Isolation for Training-Free Scene Text Editing
by: Li, Guandong, et al.
Published: (2026)
by: Li, Guandong, et al.
Published: (2026)
EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models
by: Chen, Yupeng, et al.
Published: (2024)
by: Chen, Yupeng, et al.
Published: (2024)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023)
by: Ali, Moayed Haji, et al.
Published: (2023)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
by: Motamed, Saman, et al.
Published: (2024)
by: Motamed, Saman, et al.
Published: (2024)
Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models
by: Wang, Wen, et al.
Published: (2023)
by: Wang, Wen, et al.
Published: (2023)
Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
by: Zhang, Da, et al.
Published: (2026)
by: Zhang, Da, et al.
Published: (2026)
VidCLearn: A Continual Learning Approach for Text-to-Video Generation
by: Zanchetta, Luca, et al.
Published: (2025)
by: Zanchetta, Luca, et al.
Published: (2025)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models
by: Corradini, Barbara Toniella, et al.
Published: (2024)
by: Corradini, Barbara Toniella, et al.
Published: (2024)
ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text
by: Ji, Anya, et al.
Published: (2026)
by: Ji, Anya, et al.
Published: (2026)
SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
by: Sun, Qianqian, et al.
Published: (2025)
by: Sun, Qianqian, et al.
Published: (2025)
FastEdit: Fast Text-Guided Single-Image Editing via Semantic-Aware Diffusion Fine-Tuning
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
RealCraft: Attention Control as A Tool for Zero-Shot Consistent Video Editing
by: Jin, Shutong, et al.
Published: (2023)
by: Jin, Shutong, et al.
Published: (2023)
Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
by: Liu, Kaiting, et al.
Published: (2026)
by: Liu, Kaiting, et al.
Published: (2026)
Similar Items
-
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
by: Couairon, Paul, et al.
Published: (2024) -
JAFAR: Jack up Any Feature at Any Resolution
by: Couairon, Paul, et al.
Published: (2025) -
NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering
by: Chambon, Loick, et al.
Published: (2025) -
Energy Correction Model in the Feature Space for Out-of-Distribution Detection
by: Lafon, Marc, et al.
Published: (2024) -
GalLoP: Learning Global and Local Prompts for Vision-Language Models
by: Lafon, Marc, et al.
Published: (2024)