Leveraging LLMs for On-the-Fly Instruction Guided Image Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Santos, Rodrigo, Silva, João, Branco, António |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training
by: Santos, Rodrigo, et al.
Published: (2025)
by: Santos, Rodrigo, et al.
Published: (2025)
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
by: Zhang, Kai, et al.
Published: (2023)
by: Zhang, Kai, et al.
Published: (2023)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
by: Wu, Keming, et al.
Published: (2025)
by: Wu, Keming, et al.
Published: (2025)
Bayesian Optimization for Controlled Image Editing via LLMs
by: Cai, Chengkun, et al.
Published: (2025)
by: Cai, Chengkun, et al.
Published: (2025)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
by: Yu, Shoubin, et al.
Published: (2025)
by: Yu, Shoubin, et al.
Published: (2025)
MPN: Leveraging Multilingual Patch Neuron for Cross-lingual Model Editing
by: Si, Nianwen, et al.
Published: (2024)
by: Si, Nianwen, et al.
Published: (2024)
RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives
by: Yoon, Jaehong, et al.
Published: (2024)
by: Yoon, Jaehong, et al.
Published: (2024)
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
by: Toschi, Federico, et al.
Published: (2026)
by: Toschi, Federico, et al.
Published: (2026)
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
by: Ananthram, Amith, et al.
Published: (2025)
by: Ananthram, Amith, et al.
Published: (2025)
Concept Lancet: Image Editing with Compositional Representation Transplant
by: Luo, Jinqi, et al.
Published: (2025)
by: Luo, Jinqi, et al.
Published: (2025)
Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring
by: Onsu, Murat Arda, et al.
Published: (2025)
by: Onsu, Murat Arda, et al.
Published: (2025)
MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing
by: Cheng, Liwei, et al.
Published: (2026)
by: Cheng, Liwei, et al.
Published: (2026)
MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
MANTIS: Interleaved Multi-Image Instruction Tuning
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
by: Baraldi, Lorenzo, et al.
Published: (2025)
by: Baraldi, Lorenzo, et al.
Published: (2025)
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024)
by: Hu, Hexiang, et al.
Published: (2024)
LLMs Meet Multimodal Generation and Editing: A Survey
by: He, Yingqing, et al.
Published: (2024)
by: He, Yingqing, et al.
Published: (2024)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
by: Li, Xiujun, et al.
Published: (2023)
by: Li, Xiujun, et al.
Published: (2023)
IE-Critic-R1: Advancing the Explanatory Measurement of Text-Driven Image Editing for Human Perception Alignment
by: Qu, Bowen, et al.
Published: (2025)
by: Qu, Bowen, et al.
Published: (2025)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
by: Ji, Yicheng, et al.
Published: (2025)
by: Ji, Yicheng, et al.
Published: (2025)
ETCHR: Editing To Clarify and Harness Reasoning
by: Zhang, Beichen, et al.
Published: (2026)
by: Zhang, Beichen, et al.
Published: (2026)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
by: Sarto, Sara, et al.
Published: (2025)
by: Sarto, Sara, et al.
Published: (2025)
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
by: Broomfield, Julius, et al.
Published: (2025)
by: Broomfield, Julius, et al.
Published: (2025)
DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model
by: Hong, Shibo, et al.
Published: (2026)
by: Hong, Shibo, et al.
Published: (2026)
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
by: Wada, Yuiga, et al.
Published: (2025)
by: Wada, Yuiga, et al.
Published: (2025)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
by: Cocchi, Federico, et al.
Published: (2025)
by: Cocchi, Federico, et al.
Published: (2025)
HIVE: Harnessing Human Feedback for Instructional Visual Editing
by: Zhang, Shu, et al.
Published: (2023)
by: Zhang, Shu, et al.
Published: (2023)
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
by: Huang, Han, et al.
Published: (2024)
by: Huang, Han, et al.
Published: (2024)
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
by: Jiang, Ziqi, et al.
Published: (2024)
by: Jiang, Ziqi, et al.
Published: (2024)
Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing
by: Parihar, Rishubh, et al.
Published: (2025)
by: Parihar, Rishubh, et al.
Published: (2025)
Semantic Map-based Generation of Navigation Instructions
by: Li, Chengzu, et al.
Published: (2024)
by: Li, Chengzu, et al.
Published: (2024)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
by: Yu, Zhuoran, et al.
Published: (2025)
by: Yu, Zhuoran, et al.
Published: (2025)
LLMs can Compress LLMs: Adaptive Pruning by Agents
by: Kodathala, Sai Varun, et al.
Published: (2026)
by: Kodathala, Sai Varun, et al.
Published: (2026)
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
by: Ye, Zhoutong, et al.
Published: (2025)
by: Ye, Zhoutong, et al.
Published: (2025)
Efficient Personalized Text-to-image Generation by Leveraging Textual Subspace
by: Du, Shian, et al.
Published: (2024)
by: Du, Shian, et al.
Published: (2024)
Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
by: Li, Yunxin, et al.
Published: (2023)
by: Li, Yunxin, et al.
Published: (2023)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
by: Lv, Zheqi, et al.
Published: (2025)
by: Lv, Zheqi, et al.
Published: (2025)
Similar Items
-
Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training
by: Santos, Rodrigo, et al.
Published: (2025) -
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
by: Zhang, Kai, et al.
Published: (2023) -
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
by: Wu, Keming, et al.
Published: (2025) -
Bayesian Optimization for Controlled Image Editing via LLMs
by: Cai, Chengkun, et al.
Published: (2025) -
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
by: Yu, Shoubin, et al.
Published: (2025)