POEM: Precise Object-level Editing via MLLM control
Fuente:
arXiv
Saved in:
| Main Authors: | Schouten, Marco, Kaya, Mehmet Onurcan, Belongie, Serge, Papadopoulos, Dim P. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement
by: Schouten, Marco, et al.
Published: (2026)
by: Schouten, Marco, et al.
Published: (2026)
AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
by: Lu, Kaixuan, et al.
Published: (2025)
by: Lu, Kaixuan, et al.
Published: (2025)
Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
by: Lu, Kaixuan, et al.
Published: (2025)
by: Lu, Kaixuan, et al.
Published: (2025)
Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
by: Riise, Erik, et al.
Published: (2025)
by: Riise, Erik, et al.
Published: (2025)
Efficient Test-Time Scaling for Small Vision-Language Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
prNet: Data-Driven Phase Retrieval via Stochastic Refinement
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
DDRM-PR: Fourier Phase Retrieval using Denoising Diffusion Restoration Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
I2I-PR: Deep Iterative Refinement for Phase Retrieval using Image-to-Image Diffusion Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)
Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
by: Delatolas, Thanos, et al.
Published: (2025)
by: Delatolas, Thanos, et al.
Published: (2025)
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
by: Gordon, Lucia, et al.
Published: (2026)
by: Gordon, Lucia, et al.
Published: (2026)
Visual Context-Aware Person Fall Detection
by: Nagaj, Aleksander, et al.
Published: (2024)
by: Nagaj, Aleksander, et al.
Published: (2024)
WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting
by: Wang, Lezhong, et al.
Published: (2026)
by: Wang, Lezhong, et al.
Published: (2026)
Latent Directions: A Simple Pathway to Bias Mitigation in Generative AI
by: Olmos, Carolina Lopez, et al.
Published: (2024)
by: Olmos, Carolina Lopez, et al.
Published: (2024)
Weak Cube R-CNN: Weakly Supervised 3D Detection using only 2D Bounding Boxes
by: Hansen, Andreas Lau, et al.
Published: (2025)
by: Hansen, Andreas Lau, et al.
Published: (2025)
Elysium: Exploring Object-level Perception in Videos via MLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Towards High-Quality Image Segmentation: Improving Topology Accuracy by Penalizing Neighbor Pixels
by: Valverde, Juan Miguel, et al.
Published: (2026)
by: Valverde, Juan Miguel, et al.
Published: (2026)
Labeled Data Selection for Category Discovery
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
PhysConvex: Physics-Informed 3D Dynamic Convex Radiance Fields for Reconstruction and Simulation
by: Wang, Dan, et al.
Published: (2026)
by: Wang, Dan, et al.
Published: (2026)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
by: Yu, Hong-Tao, et al.
Published: (2025)
by: Yu, Hong-Tao, et al.
Published: (2025)
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
by: Wang, Shuyu, et al.
Published: (2025)
by: Wang, Shuyu, et al.
Published: (2025)
Interaction-Consistent Object Removal via MLLM-Based Reasoning
by: Huang, Ching-Kai, et al.
Published: (2026)
by: Huang, Ching-Kai, et al.
Published: (2026)
MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
by: Kristoffersen, Oskar, et al.
Published: (2025)
by: Kristoffersen, Oskar, et al.
Published: (2025)
Moodifier: MLLM-Enhanced Emotion-Driven Image Editing
by: Ye, Jiarong, et al.
Published: (2025)
by: Ye, Jiarong, et al.
Published: (2025)
InstructX: Towards Unified Visual Editing with MLLM Guidance
by: Mou, Chong, et al.
Published: (2025)
by: Mou, Chong, et al.
Published: (2025)
Unlearning-based Neural Interpretations
by: Choi, Ching Lam, et al.
Published: (2024)
by: Choi, Ching Lam, et al.
Published: (2024)
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
by: An, Zhaochong, et al.
Published: (2025)
by: An, Zhaochong, et al.
Published: (2025)
Familiarity-Based Open-Set Recognition Under Adversarial Attacks
by: Enevoldsen, Philip, et al.
Published: (2023)
by: Enevoldsen, Philip, et al.
Published: (2023)
Interpreting Object-level Foundation Models via Visual Precision Search
by: Chen, Ruoyu, et al.
Published: (2024)
by: Chen, Ruoyu, et al.
Published: (2024)
Noise-Coded Illumination for Forensic and Photometric Video Analysis
by: Michael, Peter F., et al.
Published: (2025)
by: Michael, Peter F., et al.
Published: (2025)
Image Sculpting: Precise Object Editing with 3D Geometry Control
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
Assessing Neural Network Robustness via Adversarial Pivotal Tuning
by: Christensen, Peter Ebert, et al.
Published: (2022)
by: Christensen, Peter Ebert, et al.
Published: (2022)
Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting
by: Zhang, Xiaowen, et al.
Published: (2026)
by: Zhang, Xiaowen, et al.
Published: (2026)
SuperF: Neural Implicit Fields for Multi-Image Super-Resolution
by: Jyhne, Sander Riisøen, et al.
Published: (2025)
by: Jyhne, Sander Riisøen, et al.
Published: (2025)
Taxonomy-Aware Evaluation of Vision-Language Models
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution
by: Wang, Dan, et al.
Published: (2026)
by: Wang, Dan, et al.
Published: (2026)
Better Language Models Exhibit Higher Visual Alignment
by: Ruthardt, Jona, et al.
Published: (2024)
by: Ruthardt, Jona, et al.
Published: (2024)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2025)
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2025)
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2026)
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2026)
A$^2$-Edit: Precise Reference-Guided Image Editing of Arbitrary Objects and Ambiguous Masks
by: Zheng, Huayu, et al.
Published: (2026)
by: Zheng, Huayu, et al.
Published: (2026)
Similar Items
-
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement
by: Schouten, Marco, et al.
Published: (2026) -
AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
by: Lu, Kaixuan, et al.
Published: (2025) -
Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training
by: Lu, Kaixuan, et al.
Published: (2025) -
Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
by: Riise, Erik, et al.
Published: (2025) -
Efficient Test-Time Scaling for Small Vision-Language Models
by: Kaya, Mehmet Onurcan, et al.
Published: (2025)