Instruction-based Image Manipulation by Watching How Things Move
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Mingdeng, Zhang, Xuaner, Zheng, Yinqiang, Xia, Zhihao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rolling Shutter Correction with Intermediate Distortion Flow Estimation
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos
by: Alzayer, Hadi, et al.
Published: (2024)
by: Alzayer, Hadi, et al.
Published: (2024)
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
Restoration by Generation with Constrained Priors
by: Ding, Zheng, et al.
Published: (2023)
by: Ding, Zheng, et al.
Published: (2023)
Explorative Inbetweening of Time and Space
by: Feng, Haiwen, et al.
Published: (2024)
by: Feng, Haiwen, et al.
Published: (2024)
LEDiff: Latent Exposure Diffusion for HDR Generation
by: Wang, Chao, et al.
Published: (2024)
by: Wang, Chao, et al.
Published: (2024)
All-in-One Transferring Image Compression from Human Perception to Multi-Machine Perception
by: Zhao, Jiancheng, et al.
Published: (2025)
by: Zhao, Jiancheng, et al.
Published: (2025)
AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
by: Niu, Muyao, et al.
Published: (2025)
by: Niu, Muyao, et al.
Published: (2025)
Watch Your Steps: Local Image and Scene Editing by Text Instructions
by: Mirzaei, Ashkan, et al.
Published: (2023)
by: Mirzaei, Ashkan, et al.
Published: (2023)
ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
by: Chang, Di, et al.
Published: (2025)
by: Chang, Di, et al.
Published: (2025)
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
by: Wang, Weicheng, et al.
Published: (2026)
by: Wang, Weicheng, et al.
Published: (2026)
Motion Blur Decomposition with Cross-shutter Guidance
by: Ji, Xiang, et al.
Published: (2024)
by: Ji, Xiang, et al.
Published: (2024)
Watch Where You Move: Region-aware Dynamic Aggregation and Excitation for Gait Recognition
by: Huang, Binyuan, et al.
Published: (2025)
by: Huang, Binyuan, et al.
Published: (2025)
Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
by: Jin, Linyi, et al.
Published: (2024)
by: Jin, Linyi, et al.
Published: (2024)
KaoLRM: Repurposing Pre-trained Large Reconstruction Models for Parametric 3D Face Reconstruction
by: Zhu, Qingtian, et al.
Published: (2026)
by: Zhu, Qingtian, et al.
Published: (2026)
ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained Guidance
by: Shi, Shuwei, et al.
Published: (2024)
by: Shi, Shuwei, et al.
Published: (2024)
Fooling Polarization-based Vision using Locally Controllable Polarizing Projection
by: Li, Zhuoxiao, et al.
Published: (2023)
by: Li, Zhuoxiao, et al.
Published: (2023)
Fine-grained Defocus Blur Control for Generative Image Models
by: Shrivastava, Ayush, et al.
Published: (2025)
by: Shrivastava, Ayush, et al.
Published: (2025)
Learning to Refocus with Video Diffusion Models
by: Tedla, SaiKiran, et al.
Published: (2025)
by: Tedla, SaiKiran, et al.
Published: (2025)
ReVideo: Remake a Video with Motion and Content Control
by: Mou, Chong, et al.
Published: (2024)
by: Mou, Chong, et al.
Published: (2024)
Move and Act: Enhanced Object Manipulation and Background Integrity for Image Editing
by: Jiang, Pengfei, et al.
Published: (2024)
by: Jiang, Pengfei, et al.
Published: (2024)
Navigating Beyond Dropout: An Intriguing Solution Towards Generalizable Image Super Resolution
by: Wang, Hongjun, et al.
Published: (2024)
by: Wang, Hongjun, et al.
Published: (2024)
Generative Blocks World: Moving Things Around in Pictures
by: Vavilala, Vaibhav, et al.
Published: (2025)
by: Vavilala, Vaibhav, et al.
Published: (2025)
RS-NeRF: Neural Radiance Fields from Rolling Shutter Images
by: Niu, Muyao, et al.
Published: (2024)
by: Niu, Muyao, et al.
Published: (2024)
EventHDR: from Event to High-Speed HDR Videos and Beyond
by: Zou, Yunhao, et al.
Published: (2024)
by: Zou, Yunhao, et al.
Published: (2024)
RPBG: Towards Robust Neural Point-based Graphics in the Wild
by: Zhu, Qingtian, et al.
Published: (2024)
by: Zhu, Qingtian, et al.
Published: (2024)
MoEController: Instruction-based Arbitrary Image Manipulation with Mixture-of-Expert Controllers
by: Li, Sijia, et al.
Published: (2023)
by: Li, Sijia, et al.
Published: (2023)
HarmoQ: Harmonized Post-Training Quantization for High-Fidelity Image
by: Wang, Hongjun, et al.
Published: (2025)
by: Wang, Hongjun, et al.
Published: (2025)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
by: Niu, Muyao, et al.
Published: (2024)
by: Niu, Muyao, et al.
Published: (2024)
SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization
by: Zhuang, Peiyu, et al.
Published: (2026)
by: Zhuang, Peiyu, et al.
Published: (2026)
DreamVE: Unified Instruction-based Image and Video Editing
by: Xia, Bin, et al.
Published: (2025)
by: Xia, Bin, et al.
Published: (2025)
Not All Degradations Are Equal: A Targeted Feature Denoising Framework for Generalizable Image Super-Resolution
by: Wang, Hongjun, et al.
Published: (2025)
by: Wang, Hongjun, et al.
Published: (2025)
Latent Disentanglement for Low Light Image Enhancement
by: Zheng, Zhihao, et al.
Published: (2024)
by: Zheng, Zhihao, et al.
Published: (2024)
DensifyBeforehand: LiDAR-assisted Content-aware Densification for Efficient and Quality 3D Gaussian Splatting
by: Patt, Phurtivilai, et al.
Published: (2025)
by: Patt, Phurtivilai, et al.
Published: (2025)
EfficientHuman: Efficient Training and Reconstruction of Moving Human using Articulated 2D Gaussian
by: Tian, Hao, et al.
Published: (2025)
by: Tian, Hao, et al.
Published: (2025)
Moment-Reenacting: Inverse Motion Degradation with Cross-shutter Guidance
by: Ji, Xiang, et al.
Published: (2026)
by: Ji, Xiang, et al.
Published: (2026)
Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single Image
by: Mei, Yiqun, et al.
Published: (2024)
by: Mei, Yiqun, et al.
Published: (2024)
CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions
by: Wang, Chonghuinan, et al.
Published: (2026)
by: Wang, Chonghuinan, et al.
Published: (2026)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
by: Zhang, Huanyu, et al.
Published: (2026)
by: Zhang, Huanyu, et al.
Published: (2026)
CMFDFormer: Transformer-based Copy-Move Forgery Detection with Continual Learning
by: Liu, Yaqi, et al.
Published: (2023)
by: Liu, Yaqi, et al.
Published: (2023)
Similar Items
-
Rolling Shutter Correction with Intermediate Distortion Flow Estimation
by: Cao, Mingdeng, et al.
Published: (2024) -
Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos
by: Alzayer, Hadi, et al.
Published: (2024) -
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
by: Cao, Mingdeng, et al.
Published: (2024) -
Restoration by Generation with Constrained Priors
by: Ding, Zheng, et al.
Published: (2023) -
Explorative Inbetweening of Time and Space
by: Feng, Haiwen, et al.
Published: (2024)