From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Rajan, Anirudh Sundara, Singh, Krishna Kumar, Lee, Yong Jae |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stay-Positive: A Case for Ignoring Real Image Features in Fake Image Detection
by: Rajan, Anirudh Sundara, et al.
Published: (2025)
by: Rajan, Anirudh Sundara, et al.
Published: (2025)
Aligned Datasets Improve Detection of Latent Diffusion-Generated Images
by: Rajan, Anirudh Sundara, et al.
Published: (2024)
by: Rajan, Anirudh Sundara, et al.
Published: (2024)
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
by: Huang, Zeyi, et al.
Published: (2025)
by: Huang, Zeyi, et al.
Published: (2025)
Reasoning-Augmented Representations for Multimodal Retrieval
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Dragging with Geometry: From Pixels to Geometry-Guided Image Editing
by: Pu, Xinyu, et al.
Published: (2025)
by: Pu, Xinyu, et al.
Published: (2025)
Low-Resolution Editing is All You Need for High-Resolution Editing
by: Lee, Junsung, et al.
Published: (2025)
by: Lee, Junsung, et al.
Published: (2025)
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
by: Li, Junxian, et al.
Published: (2026)
by: Li, Junxian, et al.
Published: (2026)
RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing
by: Qu, Tianyuan, et al.
Published: (2025)
by: Qu, Tianyuan, et al.
Published: (2025)
Edit One for All: Interactive Batch Image Editing
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
IMAGAgent: Orchestrating Multi-Turn Image Editing via Constraint-Aware Planning and Reflection
by: Shen, Fei, et al.
Published: (2026)
by: Shen, Fei, et al.
Published: (2026)
Learning an Image Editing Model without Image Editing Pairs
by: Kumari, Nupur, et al.
Published: (2025)
by: Kumari, Nupur, et al.
Published: (2025)
Open-Det: An Efficient Learning Framework for Open-Ended Detection
by: Cao, Guiping, et al.
Published: (2025)
by: Cao, Guiping, et al.
Published: (2025)
I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing
by: Yu, Jinghan, et al.
Published: (2026)
by: Yu, Jinghan, et al.
Published: (2026)
Probing Visual Planning in Image Editing Models
by: Zhou, Zhimu, et al.
Published: (2026)
by: Zhou, Zhimu, et al.
Published: (2026)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
by: Li, Yuheng, et al.
Published: (2024)
by: Li, Yuheng, et al.
Published: (2024)
MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation
by: Yu, Yangcheng, et al.
Published: (2025)
by: Yu, Yangcheng, et al.
Published: (2025)
UniHuman: A Unified Model for Editing Human Images in the Wild
by: Li, Nannan, et al.
Published: (2023)
by: Li, Nannan, et al.
Published: (2023)
Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator
by: Pu, Yuhan, et al.
Published: (2026)
by: Pu, Yuhan, et al.
Published: (2026)
3D-Fixup: Advancing Photo Editing with 3D Priors
by: Cheng, Yen-Chi, et al.
Published: (2025)
by: Cheng, Yen-Chi, et al.
Published: (2025)
MediX-R1: Open Ended Medical Reinforcement Learning
by: Mullappilly, Sahal Shaji, et al.
Published: (2026)
by: Mullappilly, Sahal Shaji, et al.
Published: (2026)
From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning
by: Salgado, Alberto G. Rodriguez
Published: (2026)
by: Salgado, Alberto G. Rodriguez
Published: (2026)
LVLM-Composer's Explicit Planning for Image Generation
by: Ramsey, Spencer, et al.
Published: (2025)
by: Ramsey, Spencer, et al.
Published: (2025)
Vector Scaffolding: Inter-Scale Orchestration for Differentiable Image Vectorization
by: Lee, Jaerin, et al.
Published: (2026)
by: Lee, Jaerin, et al.
Published: (2026)
Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
by: Mo, Sicheng, et al.
Published: (2025)
by: Mo, Sicheng, et al.
Published: (2025)
From Steering to Pedalling: Do Autonomous Driving VLMs Generalize to Cyclist-Assistive Spatial Perception and Planning?
by: Nakka, Krishna Kanth, et al.
Published: (2026)
by: Nakka, Krishna Kanth, et al.
Published: (2026)
Event-based Photometric Stereo via Rotating Illumination and Per-Pixel Learning
by: Kim, Hyunwoo, et al.
Published: (2026)
by: Kim, Hyunwoo, et al.
Published: (2026)
Edit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene Editing
by: Noh, Seongrae, et al.
Published: (2026)
by: Noh, Seongrae, et al.
Published: (2026)
Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders
by: Stevens, Samuel, et al.
Published: (2025)
by: Stevens, Samuel, et al.
Published: (2025)
Generative Region-Language Pretraining for Open-Ended Object Detection
by: Lin, Chuang, et al.
Published: (2024)
by: Lin, Chuang, et al.
Published: (2024)
PhotoAgent: Agentic Photo Editing with Exploratory Visual Aesthetic Planning
by: Yao, Mingde, et al.
Published: (2026)
by: Yao, Mingde, et al.
Published: (2026)
Open-Event Procedure Planning in Instructional Videos
by: Wu, Yilu, et al.
Published: (2024)
by: Wu, Yilu, et al.
Published: (2024)
From Image- to Pixel-level: Label-efficient Hyperspectral Image Reconstruction
by: Leng, Yihong, et al.
Published: (2025)
by: Leng, Yihong, et al.
Published: (2025)
GroupDiff: Diffusion-based Group Portrait Editing
by: Jiang, Yuming, et al.
Published: (2024)
by: Jiang, Yuming, et al.
Published: (2024)
YoChameleon: Personalized Vision and Language Generation
by: Nguyen, Thao, et al.
Published: (2025)
by: Nguyen, Thao, et al.
Published: (2025)
Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation
by: Zhao, Zhonghan, et al.
Published: (2024)
by: Zhao, Zhonghan, et al.
Published: (2024)
Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels
by: Liang, Tianming, et al.
Published: (2024)
by: Liang, Tianming, et al.
Published: (2024)
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
by: Zhang, Songyan, et al.
Published: (2025)
by: Zhang, Songyan, et al.
Published: (2025)
Localized Latent Editing for Dose-Response Modeling in Botulinum Toxin Injection Planning
by: Arnaud, Estèphe, et al.
Published: (2026)
by: Arnaud, Estèphe, et al.
Published: (2026)
Similar Items
-
Stay-Positive: A Case for Ignoring Real Image Features in Fake Image Detection
by: Rajan, Anirudh Sundara, et al.
Published: (2025) -
Aligned Datasets Improve Detection of Latent Diffusion-Generated Images
by: Rajan, Anirudh Sundara, et al.
Published: (2024) -
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
by: Huang, Zeyi, et al.
Published: (2025) -
Reasoning-Augmented Representations for Multimodal Retrieval
by: Zhang, Jianrui, et al.
Published: (2026) -
Dragging with Geometry: From Pixels to Geometry-Guided Image Editing
by: Pu, Xinyu, et al.
Published: (2025)