StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Jiho, Choi, Sieun, Seo, Jaeyoon, Kim, Jihie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering
by: Park, Jiho, et al.
Published: (2026)
by: Park, Jiho, et al.
Published: (2026)
Contour-Guided Query-Based Feature Fusion for Boundary-Aware and Generalizable Cardiac Ultrasound Segmentation
by: Ullah, Zahid, et al.
Published: (2026)
by: Ullah, Zahid, et al.
Published: (2026)
Evaluating Demographic Misrepresentation in Image-to-Image Portrait Editing
by: Seo, Huichan, et al.
Published: (2026)
by: Seo, Huichan, et al.
Published: (2026)
VideoSketcher: Video Models Prior Enable Versatile Sequential Sketch Generation
by: Ren, Hui, et al.
Published: (2026)
by: Ren, Hui, et al.
Published: (2026)
DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models
by: Xing, Ximing, et al.
Published: (2023)
by: Xing, Ximing, et al.
Published: (2023)
Enhancing Creative Generation on Stable Diffusion-based Models
by: Han, Jiyeon, et al.
Published: (2025)
by: Han, Jiyeon, et al.
Published: (2025)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
by: Lee, Dosung, et al.
Published: (2025)
by: Lee, Dosung, et al.
Published: (2025)
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
by: Huang, Muye, et al.
Published: (2025)
by: Huang, Muye, et al.
Published: (2025)
Hierarchical Deep Feature Fusion and Ensemble Learning for Enhanced Brain Tumor MRI Classification
by: Ullah, Zahid, et al.
Published: (2025)
by: Ullah, Zahid, et al.
Published: (2025)
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Before We Trust Them: Decision-Making Failures in Navigation of Foundation Models
by: Han, Jua, et al.
Published: (2026)
by: Han, Jua, et al.
Published: (2026)
Enhancing Visual Re-ranking through Denoising Nearest Neighbor Graph via Continuous CRF
by: Kim, Jaeyoon, et al.
Published: (2024)
by: Kim, Jaeyoon, et al.
Published: (2024)
VisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought
by: Lee, Eunsoo, et al.
Published: (2026)
by: Lee, Eunsoo, et al.
Published: (2026)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
by: Xu, Quanxing, et al.
Published: (2026)
by: Xu, Quanxing, et al.
Published: (2026)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
by: Seo, Huichan, et al.
Published: (2025)
by: Seo, Huichan, et al.
Published: (2025)
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
by: Kwon, Minkyung, et al.
Published: (2025)
by: Kwon, Minkyung, et al.
Published: (2025)
Enhancing Generalization in Medical Visual Question Answering Tasks via Gradient-Guided Model Perturbation
by: Liu, Gang, et al.
Published: (2024)
by: Liu, Gang, et al.
Published: (2024)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
by: Song, Seokwon, et al.
Published: (2025)
by: Song, Seokwon, et al.
Published: (2025)
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
by: Kim, Wonjoong, et al.
Published: (2024)
by: Kim, Wonjoong, et al.
Published: (2024)
MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports
by: Kyung, Sunggu, et al.
Published: (2025)
by: Kyung, Sunggu, et al.
Published: (2025)
Towards Test-time Efficient Visual Place Recognition via Asymmetric Query Processing
by: Kim, Jaeyoon, et al.
Published: (2025)
by: Kim, Jaeyoon, et al.
Published: (2025)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
by: Jeong, Suchae, et al.
Published: (2025)
by: Jeong, Suchae, et al.
Published: (2025)
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
by: Lee, Seonho, et al.
Published: (2024)
by: Lee, Seonho, et al.
Published: (2024)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
by: Choi, Joonmyung, et al.
Published: (2026)
by: Choi, Joonmyung, et al.
Published: (2026)
VL-OrdinalFormer: Vision Language Guided Ordinal Transformers for Interpretable Knee Osteoarthritis Grading
by: Ullah, Zahid, et al.
Published: (2025)
by: Ullah, Zahid, et al.
Published: (2025)
Exploring Kolmogorov-Arnold Network Expansions in Vision Transformers for Mitigating Catastrophic Forgetting in Continual Learning
by: Ullah, Zahid, et al.
Published: (2025)
by: Ullah, Zahid, et al.
Published: (2025)
DAM-Seg: Anatomically accurate cardiac segmentation using Dense Associative Networks
by: Ullah, Zahid, et al.
Published: (2025)
by: Ullah, Zahid, et al.
Published: (2025)
Hybrid Ensemble Approaches: Optimal Deep Feature Fusion and Hyperparameter-Tuned Classifier Ensembling for Enhanced Brain Tumor Classification
by: Ullah, Zahid, et al.
Published: (2025)
by: Ullah, Zahid, et al.
Published: (2025)
Advancing Medical Image Segmentation: Morphology-Driven Learning with Diffusion Transformer
by: Kang, Sungmin, et al.
Published: (2024)
by: Kang, Sungmin, et al.
Published: (2024)
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
by: Movva, Prahitha, et al.
Published: (2025)
by: Movva, Prahitha, et al.
Published: (2025)
Generalizing Visual Question Answering from Synthetic to Human-Written Questions via a Chain of QA with a Large Language Model
by: Kim, Taehee, et al.
Published: (2024)
by: Kim, Taehee, et al.
Published: (2024)
Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
by: Xue, Junxiao, et al.
Published: (2024)
by: Xue, Junxiao, et al.
Published: (2024)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
by: Lim, Su Hyeon, et al.
Published: (2024)
by: Lim, Su Hyeon, et al.
Published: (2024)
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)
by: Rosenfeld, Amir, et al.
Published: (2025)
SwiftSketch: A Diffusion Model for Image-to-Vector Sketch Generation
by: Arar, Ellie, et al.
Published: (2025)
by: Arar, Ellie, et al.
Published: (2025)
Towards Flexible Evaluation for Generative Visual Question Answering
by: Ji, Huishan, et al.
Published: (2024)
by: Ji, Huishan, et al.
Published: (2024)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
by: Choi, Jiho, et al.
Published: (2026)
by: Choi, Jiho, et al.
Published: (2026)
Similar Items
-
SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering
by: Park, Jiho, et al.
Published: (2026) -
Contour-Guided Query-Based Feature Fusion for Boundary-Aware and Generalizable Cardiac Ultrasound Segmentation
by: Ullah, Zahid, et al.
Published: (2026) -
Evaluating Demographic Misrepresentation in Image-to-Image Portrait Editing
by: Seo, Huichan, et al.
Published: (2026) -
VideoSketcher: Video Models Prior Enable Versatile Sequential Sketch Generation
by: Ren, Hui, et al.
Published: (2026) -
DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models
by: Xing, Ximing, et al.
Published: (2023)