Point and Instruct: Enabling Precise Image Editing by Unifying Direct Manipulation and Text Instructions
Fuente:
arXiv
Guardado en:
| Autores principales: | Helbling, Alec, Lee, Seongmin, Chau, Polo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ClickDiffusion: Harnessing LLMs for Interactive Precise Image Editing
por: Helbling, Alec, et al.
Publicado: (2024)
por: Helbling, Alec, et al.
Publicado: (2024)
Mobile Fitting Room: On-device Virtual Try-on via Diffusion Models
por: Blalock, Justin, et al.
Publicado: (2024)
por: Blalock, Justin, et al.
Publicado: (2024)
Transformer Explainer: Interactive Learning of Text-Generative Models
por: Cho, Aeree, et al.
Publicado: (2024)
por: Cho, Aeree, et al.
Publicado: (2024)
LLM Attributor: Interactive Visual Attribution for LLM Generation
por: Lee, Seongmin, et al.
Publicado: (2024)
por: Lee, Seongmin, et al.
Publicado: (2024)
InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMs
por: Zhou, Zhongyi, et al.
Publicado: (2023)
por: Zhou, Zhongyi, et al.
Publicado: (2023)
Effective Guidance for Model Attention with Simple Yes-no Annotations
por: Lee, Seongmin, et al.
Publicado: (2024)
por: Lee, Seongmin, et al.
Publicado: (2024)
Interactive Visual Learning for Stable Diffusion
por: Lee, Seongmin, et al.
Publicado: (2024)
por: Lee, Seongmin, et al.
Publicado: (2024)
InstructEdit: Instruction-based Knowledge Editing for Large Language Models
por: Zhang, Ningyu, et al.
Publicado: (2024)
por: Zhang, Ningyu, et al.
Publicado: (2024)
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
por: Kyaw, Alexander Htet, et al.
Publicado: (2025)
por: Kyaw, Alexander Htet, et al.
Publicado: (2025)
AVIN-Chat: An Audio-Visual Interactive Chatbot System with Emotional State Tuning
por: Park, Chanhyuk, et al.
Publicado: (2024)
por: Park, Chanhyuk, et al.
Publicado: (2024)
Sparse Activation Editing for Reliable Instruction Following in Narratives
por: Zhao, Runcong, et al.
Publicado: (2025)
por: Zhao, Runcong, et al.
Publicado: (2025)
TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones
por: Tu, Minghao, et al.
Publicado: (2025)
por: Tu, Minghao, et al.
Publicado: (2025)
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
por: Lee, Seongmin, et al.
Publicado: (2023)
por: Lee, Seongmin, et al.
Publicado: (2023)
Mapping the Mind of an Instruction-based Image Editing using SMILE
por: Dehghani, Zeinab, et al.
Publicado: (2024)
por: Dehghani, Zeinab, et al.
Publicado: (2024)
HelpViz: Automatic Generation of Contextual Visual MobileTutorials from Text-Based Instructions
por: Zhong, Mingyuan, et al.
Publicado: (2021)
por: Zhong, Mingyuan, et al.
Publicado: (2021)
Autoware.Flex: Human-Instructed Dynamically Reconfigurable Autonomous Driving Systems
por: Song, Ziwei, et al.
Publicado: (2024)
por: Song, Ziwei, et al.
Publicado: (2024)
Synthetic Human Memories: AI-Edited Images and Videos Can Implant False Memories and Distort Recollection
por: Pataranutaporn, Pat, et al.
Publicado: (2024)
por: Pataranutaporn, Pat, et al.
Publicado: (2024)
Logic-Scaffolding: Personalized Aspect-Instructed Recommendation Explanation Generation using LLMs
por: Rahdari, Behnam, et al.
Publicado: (2023)
por: Rahdari, Behnam, et al.
Publicado: (2023)
LEDITS++: Limitless Image Editing using Text-to-Image Models
por: Brack, Manuel, et al.
Publicado: (2023)
por: Brack, Manuel, et al.
Publicado: (2023)
PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement
por: Wang, Zhijie, et al.
Publicado: (2024)
por: Wang, Zhijie, et al.
Publicado: (2024)
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models
por: Ou, Yixin, et al.
Publicado: (2024)
por: Ou, Yixin, et al.
Publicado: (2024)
ExpressEdit: Video Editing with Natural Language and Sketching
por: Tilekbay, Bekzat, et al.
Publicado: (2024)
por: Tilekbay, Bekzat, et al.
Publicado: (2024)
EditScribe: Non-Visual Image Editing with Natural Language Verification Loops
por: Chang, Ruei-Che, et al.
Publicado: (2024)
por: Chang, Ruei-Che, et al.
Publicado: (2024)
AgentCTG: Harnessing Multi-Agent Collaboration for Fine-Grained Precise Control in Text Generation
por: Zhou, Xinxu, et al.
Publicado: (2025)
por: Zhou, Xinxu, et al.
Publicado: (2025)
Insights Informed Generative AI for Design: Incorporating Real-world Data for Text-to-Image Output
por: Gupta, Richa, et al.
Publicado: (2025)
por: Gupta, Richa, et al.
Publicado: (2025)
Enabling Multi-Agent Systems as Learning Designers: Applying Learning Sciences to AI Instructional Design
por: Wang, Jiayi, et al.
Publicado: (2025)
por: Wang, Jiayi, et al.
Publicado: (2025)
What's Next? Exploring Utilization, Challenges, and Future Directions of AI-Generated Image Tools in Graphic Design
por: Tang, Yuying, et al.
Publicado: (2024)
por: Tang, Yuying, et al.
Publicado: (2024)
Goetterfunke: Creativity in Machinae Sapiens. About the Qualitative Shift in Generative AI with a Focus on Text-To-Image
por: Knappe, Jens
Publicado: (2024)
por: Knappe, Jens
Publicado: (2024)
ChatHouseDiffusion: Prompt-Guided Generation and Editing of Floor Plans
por: Qin, Sizhong, et al.
Publicado: (2024)
por: Qin, Sizhong, et al.
Publicado: (2024)
Aria-UI: Visual Grounding for GUI Instructions
por: Yang, Yuhao, et al.
Publicado: (2024)
por: Yang, Yuhao, et al.
Publicado: (2024)
Lost in Instructions: Study of Blind Users' Experiences with DIY Manuals and AI-Rewritten Instructions for Assembly, Operation, and Troubleshooting of Tangible Products
por: Reddy, Monalika Padma, et al.
Publicado: (2026)
por: Reddy, Monalika Padma, et al.
Publicado: (2026)
Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry Practitioners
por: Wang, Qiaosi, et al.
Publicado: (2026)
por: Wang, Qiaosi, et al.
Publicado: (2026)
Beyond Instructed Tasks: Recognizing In-the-Wild Reading Behaviors in the Classroom Using Eye Tracking
por: Davalos, Eduardo, et al.
Publicado: (2025)
por: Davalos, Eduardo, et al.
Publicado: (2025)
Embedding Democratic Values into Social Media AIs via Societal Objective Functions
por: Jia, Chenyan, et al.
Publicado: (2023)
por: Jia, Chenyan, et al.
Publicado: (2023)
Rewriting Conversational Utterances with Instructed Large Language Models
por: Galimzhanova, Elnara, et al.
Publicado: (2024)
por: Galimzhanova, Elnara, et al.
Publicado: (2024)
KEditVis: A Visual Analytics System for Knowledge Editing of Large Language Models
por: Chen, Zhenning, et al.
Publicado: (2026)
por: Chen, Zhenning, et al.
Publicado: (2026)
A Directed Graph Model and Experimental Framework for Design and Study of Time-Dependent Text Visualisation
por: Fan, Songhai, et al.
Publicado: (2026)
por: Fan, Songhai, et al.
Publicado: (2026)
Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt Engineering
por: Zhang, Tanghaoran, et al.
Publicado: (2024)
por: Zhang, Tanghaoran, et al.
Publicado: (2024)
Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South
por: Rastogi, Charvi, et al.
Publicado: (2026)
por: Rastogi, Charvi, et al.
Publicado: (2026)
AI Meets Maritime Training: Precision Analytics for Enhanced Safety and Performance
por: Lall, Vishakha, et al.
Publicado: (2025)
por: Lall, Vishakha, et al.
Publicado: (2025)
Ejemplares similares
-
ClickDiffusion: Harnessing LLMs for Interactive Precise Image Editing
por: Helbling, Alec, et al.
Publicado: (2024) -
Mobile Fitting Room: On-device Virtual Try-on via Diffusion Models
por: Blalock, Justin, et al.
Publicado: (2024) -
Transformer Explainer: Interactive Learning of Text-Generative Models
por: Cho, Aeree, et al.
Publicado: (2024) -
LLM Attributor: Interactive Visual Attribution for LLM Generation
por: Lee, Seongmin, et al.
Publicado: (2024) -
InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMs
por: Zhou, Zhongyi, et al.
Publicado: (2023)