Learning Complex Non-Rigid Image Edits from Multimodal Conditioning
Fuente:
arXiv
Saved in:
| Main Authors: | Warner, Nikolai, Kolb, Jack, Hahn, Meera, Birodkar, Vighnesh, Huang, Jonathan, Essa, Irfan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Learning of Disentangled Representations from Video
by: Denton, Remi, et al.
Published: (2017)
by: Denton, Remi, et al.
Published: (2017)
MoCHA: Denoising Caption Supervision for Motion-Text Retrieval
by: Warner, Nikolai, et al.
Published: (2026)
by: Warner, Nikolai, et al.
Published: (2026)
AugLift: Depth-Aware Input Reparameterization Improves Domain Generalization in 2D-to-3D Pose Lifting
by: Warner, Nikolai, et al.
Published: (2025)
by: Warner, Nikolai, et al.
Published: (2025)
MALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation
by: Yu, Sihyun, et al.
Published: (2025)
by: Yu, Sihyun, et al.
Published: (2025)
Sample what you cant compress
by: Birodkar, Vighnesh, et al.
Published: (2024)
by: Birodkar, Vighnesh, et al.
Published: (2024)
Leveraging Procedural Knowledge and Task Hierarchies for Efficient Instructional Video Pre-training
by: Samel, Karan, et al.
Published: (2025)
by: Samel, Karan, et al.
Published: (2025)
SLAIM: Robust Dense Neural SLAM for Online Tracking and Mapping
by: Cartillier, Vincent, et al.
Published: (2024)
by: Cartillier, Vincent, et al.
Published: (2024)
3D Semantic MapNet: Building Maps for Multi-Object Re-Identification in 3D
by: Cartillier, Vincent, et al.
Published: (2024)
by: Cartillier, Vincent, et al.
Published: (2024)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
by: Marmon, Andrew, et al.
Published: (2024)
by: Marmon, Andrew, et al.
Published: (2024)
FlexiEdit: Frequency-Aware Latent Refinement for Enhanced Non-Rigid Editing
by: Koo, Gwanhyeong, et al.
Published: (2024)
by: Koo, Gwanhyeong, et al.
Published: (2024)
HierSum: A Global and Local Attention Mechanism for Video Summarization
by: Beedu, Apoorva, et al.
Published: (2025)
by: Beedu, Apoorva, et al.
Published: (2025)
Mamba Fusion: Learning Actions Through Questioning
by: Dong, Zhikang, et al.
Published: (2024)
by: Dong, Zhikang, et al.
Published: (2024)
Image-Based Classification of Olive Varieties Native to Turkiye Using Multiple Deep Learning Architectures: Analysis of Performance, Complexity, and Generalization
by: Karatas, Hatice, et al.
Published: (2026)
by: Karatas, Hatice, et al.
Published: (2026)
Unified Diffusion-Based Rigid and Non-Rigid Editing with Text and Image Guidance
by: Wang, Jiacheng, et al.
Published: (2024)
by: Wang, Jiacheng, et al.
Published: (2024)
CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editing
by: Wang, Yucheng, et al.
Published: (2026)
by: Wang, Yucheng, et al.
Published: (2026)
ReEdit: Multimodal Exemplar-Based Image Editing with Diffusion Models
by: Srivastava, Ashutosh, et al.
Published: (2024)
by: Srivastava, Ashutosh, et al.
Published: (2024)
Robust Rigid and Non-Rigid Medical Image Registration Using Learnable Edge Kernels
by: Siyal, Ahsan Raza, et al.
Published: (2025)
by: Siyal, Ahsan Raza, et al.
Published: (2025)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
by: Yu, Lijun, et al.
Published: (2023)
by: Yu, Lijun, et al.
Published: (2023)
Exploring Efficient Foundational Multi-modal Models for Video Summarization
by: Samel, Karan, et al.
Published: (2024)
by: Samel, Karan, et al.
Published: (2024)
EditCLIP: Representation Learning for Image Editing
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies
by: Wang, Chenglin, et al.
Published: (2025)
by: Wang, Chenglin, et al.
Published: (2025)
SeedEdit: Align Image Re-Generation to Image Editing
by: Shi, Yichun, et al.
Published: (2024)
by: Shi, Yichun, et al.
Published: (2024)
Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
EditScout: Locating Forged Regions from Diffusion-based Edited Images with Multimodal LLM
by: Nguyen, Quang, et al.
Published: (2024)
by: Nguyen, Quang, et al.
Published: (2024)
ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
by: Chang, Di, et al.
Published: (2025)
by: Chang, Di, et al.
Published: (2025)
VideoPoet: A Large Language Model for Zero-Shot Video Generation
by: Kondratyuk, Dan, et al.
Published: (2023)
by: Kondratyuk, Dan, et al.
Published: (2023)
Edit360: 2D Image Edits to 3D Assets from Any Angle
by: Huang, Junchao, et al.
Published: (2025)
by: Huang, Junchao, et al.
Published: (2025)
X-RAFT: Cross-Modal Non-Rigid Registration of Blue and White Light Neurosurgical Hyperspectral Images
by: Budd, Charlie, et al.
Published: (2025)
by: Budd, Charlie, et al.
Published: (2025)
Beyond Rigid: Benchmarking Non-Rigid Video Editing
by: Qu, Bingzheng, et al.
Published: (2026)
by: Qu, Bingzheng, et al.
Published: (2026)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
FunEditor: Achieving Complex Image Edits via Function Aggregation with Diffusion Models
by: Samadi, Mohammadreza, et al.
Published: (2024)
by: Samadi, Mohammadreza, et al.
Published: (2024)
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
by: Xu, Haohang, et al.
Published: (2026)
by: Xu, Haohang, et al.
Published: (2026)
Learning Where to Edit Vision Transformers
by: Yang, Yunqiao, et al.
Published: (2024)
by: Yang, Yunqiao, et al.
Published: (2024)
ParallelEdits: Efficient Multi-object Image Editing
by: Huang, Mingzhen, et al.
Published: (2024)
by: Huang, Mingzhen, et al.
Published: (2024)
EmoEdit: Evoking Emotions through Image Manipulation
by: Yang, Jingyuan, et al.
Published: (2024)
by: Yang, Jingyuan, et al.
Published: (2024)
ATR-UMMIM: A Benchmark Dataset for UAV-Based Multimodal Image Registration under Complex Imaging Conditions
by: Bin, Kangcheng, et al.
Published: (2025)
by: Bin, Kangcheng, et al.
Published: (2025)
EditAR: Unified Conditional Generation with Autoregressive Models
by: Mu, Jiteng, et al.
Published: (2025)
by: Mu, Jiteng, et al.
Published: (2025)
EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics
by: Nguyen, Van-Loc, et al.
Published: (2026)
by: Nguyen, Van-Loc, et al.
Published: (2026)
MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching
by: Huang, Jiahui, et al.
Published: (2026)
by: Huang, Jiahui, et al.
Published: (2026)
Self-Supervised Learning for Multimodal Non-Rigid 3D Shape Matching
by: Cao, Dongliang, et al.
Published: (2023)
by: Cao, Dongliang, et al.
Published: (2023)
Similar Items
-
Unsupervised Learning of Disentangled Representations from Video
by: Denton, Remi, et al.
Published: (2017) -
MoCHA: Denoising Caption Supervision for Motion-Text Retrieval
by: Warner, Nikolai, et al.
Published: (2026) -
AugLift: Depth-Aware Input Reparameterization Improves Domain Generalization in 2D-to-3D Pose Lifting
by: Warner, Nikolai, et al.
Published: (2025) -
MALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation
by: Yu, Sihyun, et al.
Published: (2025) -
Sample what you cant compress
by: Birodkar, Vighnesh, et al.
Published: (2024)