Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Weifeng, Wei, Xinyu, An, Ruichuan, Gao, Peng, Zou, Bocheng, Luo, Yulin, Huang, Siyuan, Zhang, Shanghang, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
by: Luo, Yulin, et al.
Published: (2024)
by: Luo, Yulin, et al.
Published: (2024)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
by: Zhang, Qizhe, et al.
Published: (2023)
by: Zhang, Qizhe, et al.
Published: (2023)
Agent Skills Should Go Beyond Text: The Case for Visual Skills
by: Xu, Binxiao, et al.
Published: (2026)
by: Xu, Binxiao, et al.
Published: (2026)
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026)
by: Lin, Weifeng, et al.
Published: (2026)
Vision Language Models See What You Want but not What You See
by: Gao, Qingying, et al.
Published: (2024)
by: Gao, Qingying, et al.
Published: (2024)
Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
by: Hu, Hongsheng, et al.
Published: (2024)
by: Hu, Hongsheng, et al.
Published: (2024)
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
by: Xiao, Junhao, et al.
Published: (2026)
by: Xiao, Junhao, et al.
Published: (2026)
Spatially Selective Imaging in Color: What You See is What You Want
by: John You En Chan, et al.
Published: (2024)
by: John You En Chan, et al.
Published: (2024)
PixWizard: Versatile Image-to-Image Visual Assistant with Open-Language Instructions
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
by: Dai, Gaole, et al.
Published: (2025)
by: Dai, Gaole, et al.
Published: (2025)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
by: Fang, Bo, et al.
Published: (2025)
by: Fang, Bo, et al.
Published: (2025)
NeIn: Telling What You Don't Want
by: Bui, Nhat-Tan, et al.
Published: (2024)
by: Bui, Nhat-Tan, et al.
Published: (2024)
UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
by: An, Ruichuan, et al.
Published: (2025)
by: An, Ruichuan, et al.
Published: (2025)
Tell Me What You Want (What You Really, Really Want): Addressing the Expectation Gap for Goal Conveyance from Humans to Robots
by: Leahy, Kevin, et al.
Published: (2024)
by: Leahy, Kevin, et al.
Published: (2024)
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025)
by: Lin, Weifeng, et al.
Published: (2025)
You Can't Always Get What You Want: Games of Ordered Preference
by: Lee, Dong Ho, et al.
Published: (2024)
by: Lee, Dong Ho, et al.
Published: (2024)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
by: Chen, Xinyan, et al.
Published: (2025)
by: Chen, Xinyan, et al.
Published: (2025)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
What You Want the Future to Be... Children's Services and Library Administrators
by: Summers, William F.
Published: (1977)
by: Summers, William F.
Published: (1977)
Visual Room 2.0: Seeing is Not Understanding for MLLMs
by: Li, Haokun, et al.
Published: (2025)
by: Li, Haokun, et al.
Published: (2025)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
Interactive3D: Create What You Want by Interactive 3D Generation
by: Dong, Shaocong, et al.
Published: (2024)
by: Dong, Shaocong, et al.
Published: (2024)
Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
by: Choi, Yumin, et al.
Published: (2025)
by: Choi, Yumin, et al.
Published: (2025)
What You Prompt is What You Get: Increasing Transparency of Prompting Using Prompt Cards
by: Caut, Amandine M., et al.
Published: (2026)
by: Caut, Amandine M., et al.
Published: (2026)
Grasp What You Want: Embodied Dexterous Grasping System Driven by Your Voice
by: Li, Junliang, et al.
Published: (2024)
by: Li, Junliang, et al.
Published: (2024)
Count What You Want: Exemplar Identification and Few-shot Counting of Human Actions in the Wild
by: Huang, Yifeng, et al.
Published: (2023)
by: Huang, Yifeng, et al.
Published: (2023)
What Teens Want: Thirty Graphic Novels You Can't Live Without.
by: Gorman, Michele
Published: (2002)
by: Gorman, Michele
Published: (2002)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
Do You Want to Hang Out? Understanding the Positive and Negative Consequences of Receiving Social Activity Invitations at Work
by: Chieh‐Yu (Joy) Lin, et al.
Published: (2025)
by: Chieh‐Yu (Joy) Lin, et al.
Published: (2025)
Comprehending Columbine
by: Larkin, Ralph
Published: (2017)
by: Larkin, Ralph
Published: (2017)
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
by: Li, Yian, et al.
Published: (2024)
by: Li, Yian, et al.
Published: (2024)
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs
by: Han, Yujin, et al.
Published: (2025)
by: Han, Yujin, et al.
Published: (2025)
What Boomers Want
by: Dempsey, Beth
Published: (2007)
by: Dempsey, Beth
Published: (2007)
What Kids Want.
by: Gallo, Donald R.
Published: (1995)
by: Gallo, Donald R.
Published: (1995)
YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information
by: Wang, Chien-Yao, et al.
Published: (2024)
by: Wang, Chien-Yao, et al.
Published: (2024)
What You Always Wanted to Know about the Card Catalog and Were Afraid to Ask!
by: DuPree, Sherry Sherrod
Published: (1978)
by: DuPree, Sherry Sherrod
Published: (1978)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
by: Lin, Junming, et al.
Published: (2024)
by: Lin, Junming, et al.
Published: (2024)
Similar Items
-
LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
by: Luo, Yulin, et al.
Published: (2024) -
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
by: Zhang, Qizhe, et al.
Published: (2023) -
Agent Skills Should Go Beyond Text: The Case for Visual Skills
by: Xu, Binxiao, et al.
Published: (2026) -
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026) -
Vision Language Models See What You Want but not What You See
by: Gao, Qingying, et al.
Published: (2024)