SketchVLM: Vision language models can annotate images to explain thoughts and guide users
Fuente:
arXiv
Saved in:
| Main Authors: | Collins, Brandon, Bolton, Logan, Nguyen, Hung Huy, Taesiri, Mohammad Reza, Bui, Trung, Nguyen, Anh Totti |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025)
by: Nguyen, Tin, et al.
Published: (2025)
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
anguyen8/vision-llms-are-blind: official
by: Pooyan R, et al.
Published: (2026)
by: Pooyan R, et al.
Published: (2026)
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)
by: Giang, et al.
Published: (2023)
Vision Language Models are Biased
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
PageGuide: Browser extension to assist users in navigating a webpage and locating information
by: Nguyen, Tin, et al.
Published: (2026)
by: Nguyen, Tin, et al.
Published: (2026)
Allowing humans to interactively guide machines where to look does not always improve human-AI team's classification accuracy
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
by: Nguyen, Hung Huy, et al.
Published: (2025)
by: Nguyen, Hung Huy, et al.
Published: (2025)
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
by: Pham, Thang M., et al.
Published: (2024)
by: Pham, Thang M., et al.
Published: (2024)
GlitchBench: Can large multimodal models detect video game glitches?
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
Leveraging Habitat Information for Fine-grained Bird Identification
by: Nguyen, Tin, et al.
Published: (2023)
by: Nguyen, Tin, et al.
Published: (2023)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM
by: Dao, Alan, et al.
Published: (2025)
by: Dao, Alan, et al.
Published: (2025)
Explaining Graph Neural Networks via Structure-aware Interaction Index
by: Bui, Ngoc, et al.
Published: (2024)
by: Bui, Ngoc, et al.
Published: (2024)
VideoGameBunny: Towards vision assistants for video games
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
LiteGPT: Large Vision-Language Model for Joint Chest X-ray Localization and Classification Task
by: Le-Duc, Khai, et al.
Published: (2024)
by: Le-Duc, Khai, et al.
Published: (2024)
Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
by: Zhou, Runtao, et al.
Published: (2025)
by: Zhou, Runtao, et al.
Published: (2025)
How well can a large language model explain business processes as perceived by users?
by: Fahland, Dirk, et al.
Published: (2024)
by: Fahland, Dirk, et al.
Published: (2024)
Primordial deuterium abundance from calculations of $p(n,γ)$ and $d(p,γ)$ reactions within potential-model approach
by: Anh, Nguyen Le, et al.
Published: (2026)
by: Anh, Nguyen Le, et al.
Published: (2026)
VisionGuard: Synergistic Framework for Helmet Violation Detection
by: Nguyen, Lam-Huy, et al.
Published: (2025)
by: Nguyen, Lam-Huy, et al.
Published: (2025)
Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models
by: Nguyen, Minh Khoi, et al.
Published: (2026)
by: Nguyen, Minh Khoi, et al.
Published: (2026)
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2023)
by: Nguyen, Huy, et al.
Published: (2023)
Fourier-Attentive Representation Learning: A Fourier-Guided Framework for Few-Shot Generalization in Vision-Language Models
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
by: Pham, Hieu Dinh Trung, et al.
Published: (2025)
Structured Pruning for Diverse Best-of-N Reasoning Optimization
by: Nguyen, Hieu Trung, et al.
Published: (2025)
by: Nguyen, Hieu Trung, et al.
Published: (2025)
RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
by: Park, Junghyun, et al.
Published: (2025)
by: Park, Junghyun, et al.
Published: (2025)
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
by: Nguyen, Huu Tien, et al.
Published: (2025)
by: Nguyen, Huu Tien, et al.
Published: (2025)
ViMQ: A Vietnamese Medical Question Dataset for Healthcare Dialogue System Development
by: Huy, Ta Duc, et al.
Published: (2023)
by: Huy, Ta Duc, et al.
Published: (2023)
LoG-VMamba: Local-Global Vision Mamba for Medical Image Segmentation
by: Dang, Trung Dinh Quoc, et al.
Published: (2024)
by: Dang, Trung Dinh Quoc, et al.
Published: (2024)
XMainframe: A Large Language Model for Mainframe Modernization
by: Dau, Anh T. V., et al.
Published: (2024)
by: Dau, Anh T. V., et al.
Published: (2024)
CT to PET Translation: A Large-scale Dataset and Domain-Knowledge-Guided Diffusion Approach
by: Nguyen, Dac Thai, et al.
Published: (2024)
by: Nguyen, Dac Thai, et al.
Published: (2024)
EnseSmells: Deep ensemble and programming language models for automated code smells detection
by: Ho, Anh, et al.
Published: (2025)
by: Ho, Anh, et al.
Published: (2025)
Q-learning-based Opportunistic Communication for Real-time Mobile Air Quality Monitoring Systems
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
Larger Is Not Always Better: Leveraging Structured Code Diffs for Comment Inconsistency Detection
by: Nguyen, Phong, et al.
Published: (2025)
by: Nguyen, Phong, et al.
Published: (2025)
Interpretable LLM-based Table Question Answering
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
Generative Conditional Distributions by Neural (Entropic) Optimal Transport
by: Nguyen, Bao, et al.
Published: (2024)
by: Nguyen, Bao, et al.
Published: (2024)
Task-driven Layerwise Additive Activation Intervention
by: Nguyen, Hieu Trung, et al.
Published: (2025)
by: Nguyen, Hieu Trung, et al.
Published: (2025)
VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
Similar Items
-
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024) -
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025) -
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
by: Taesiri, Mohammad Reza, et al.
Published: (2025) -
anguyen8/vision-llms-are-blind: official
by: Pooyan R, et al.
Published: (2026) -
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025)