See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
Fuente:
arXiv
Guardado en:
| Autores principales: | Park, Jaehyun, Ahn, Minyoung, Kim, Minkyu, Lee, Jonghyun, Lee, Jae-Gil, Park, Dongmin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
por: Kim, Minkyu, et al.
Publicado: (2026)
por: Kim, Minkyu, et al.
Publicado: (2026)
Refining Visual Artifacts in Diffusion Models via Explainable AI-based Flaw Activation Maps
por: Lee, Seoyeon, et al.
Publicado: (2025)
por: Lee, Seoyeon, et al.
Publicado: (2025)
Test-time Alignment of Diffusion Models without Reward Over-optimization
por: Kim, Sunwoo, et al.
Publicado: (2025)
por: Kim, Sunwoo, et al.
Publicado: (2025)
Active Learning for Continual Learning: Keeping the Past Alive in the Present
por: Park, Jaehyun, et al.
Publicado: (2025)
por: Park, Jaehyun, et al.
Publicado: (2025)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
por: Park, Dongmin, et al.
Publicado: (2024)
por: Park, Dongmin, et al.
Publicado: (2024)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
por: Ahn, Young Jin, et al.
Publicado: (2024)
por: Ahn, Young Jin, et al.
Publicado: (2024)
Impact of Regularization on Calibration and Robustness: from the Representation Space Perspective
por: Park, Jonghyun, et al.
Publicado: (2024)
por: Park, Jonghyun, et al.
Publicado: (2024)
TRACE: Your Diffusion Model is Secretly an Instance Edge Detector
por: Jo, Sanghyun, et al.
Publicado: (2025)
por: Jo, Sanghyun, et al.
Publicado: (2025)
DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs
por: Park, Minyoung, et al.
Publicado: (2026)
por: Park, Minyoung, et al.
Publicado: (2026)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
por: An, Sojung, et al.
Publicado: (2025)
por: An, Sojung, et al.
Publicado: (2025)
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
por: Park, Jonghyun, et al.
Publicado: (2025)
por: Park, Jonghyun, et al.
Publicado: (2025)
Text-Aware Image Restoration with Diffusion Models
por: Min, Jaewon, et al.
Publicado: (2025)
por: Min, Jaewon, et al.
Publicado: (2025)
Accelerating Diffusion via Hybrid Data-Pipeline Parallelism Based on Conditional Guidance Scheduling
por: Jung, Euisoo, et al.
Publicado: (2026)
por: Jung, Euisoo, et al.
Publicado: (2026)
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
por: Hur, Jiwan, et al.
Publicado: (2024)
por: Hur, Jiwan, et al.
Publicado: (2024)
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task
por: Kim, Yunho, et al.
Publicado: (2024)
por: Kim, Yunho, et al.
Publicado: (2024)
Modeling Stereo-Confidence Out of the End-to-End Stereo-Matching Network via Disparity Plane Sweep
por: Lee, Jae Young, et al.
Publicado: (2024)
por: Lee, Jae Young, et al.
Publicado: (2024)
Compose and Conquer: Diffusion-Based 3D Depth Aware Composable Image Synthesis
por: Lee, Jonghyun, et al.
Publicado: (2024)
por: Lee, Jonghyun, et al.
Publicado: (2024)
DataFreeShield: Defending Adversarial Attacks without Training Data
por: Lee, Hyeyoon, et al.
Publicado: (2024)
por: Lee, Hyeyoon, et al.
Publicado: (2024)
H2O-SDF: Two-phase Learning for 3D Indoor Reconstruction using Object Surface Fields
por: Park, Minyoung, et al.
Publicado: (2024)
por: Park, Minyoung, et al.
Publicado: (2024)
Mitigating Cross-Image Information Leakage in LVLMs for Multi-Image Tasks
por: Park, Yeji, et al.
Publicado: (2025)
por: Park, Yeji, et al.
Publicado: (2025)
Discovering and Mitigating Visual Biases through Keyword Explanation
por: Kim, Younghyun, et al.
Publicado: (2023)
por: Kim, Younghyun, et al.
Publicado: (2023)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
por: Choi, Kanghyun, et al.
Publicado: (2024)
por: Choi, Kanghyun, et al.
Publicado: (2024)
Active Prompt Learning in Vision Language Models
por: Bang, Jihwan, et al.
Publicado: (2023)
por: Bang, Jihwan, et al.
Publicado: (2023)
Diffusion-based Data Augmentation and Knowledge Distillation with Generated Soft Labels Solving Data Scarcity Problems of SAR Oil Spill Segmentation
por: Moon, Jaeho, et al.
Publicado: (2024)
por: Moon, Jaeho, et al.
Publicado: (2024)
Probing Visual Language Priors in VLMs
por: Luo, Tiange, et al.
Publicado: (2024)
por: Luo, Tiange, et al.
Publicado: (2024)
Cardiac Segmentation on CT Images through Shape-Aware Contour Attentions
por: Park, Sanguk, et al.
Publicado: (2021)
por: Park, Sanguk, et al.
Publicado: (2021)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
por: Lee, Sanghyun, et al.
Publicado: (2025)
por: Lee, Sanghyun, et al.
Publicado: (2025)
PointFix: Learning to Fix Domain Bias for Robust Online Stereo Adaptation
por: Kim, Kwonyoung, et al.
Publicado: (2022)
por: Kim, Kwonyoung, et al.
Publicado: (2022)
Stereo-Matching Knowledge Distilled Monocular Depth Estimation Filtered by Multiple Disparity Consistency
por: Ka, Woonghyun, et al.
Publicado: (2024)
por: Ka, Woonghyun, et al.
Publicado: (2024)
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
por: Hong, Rui, et al.
Publicado: (2026)
por: Hong, Rui, et al.
Publicado: (2026)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
por: Lin, Weifeng, et al.
Publicado: (2024)
por: Lin, Weifeng, et al.
Publicado: (2024)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
por: Han, Woojung, et al.
Publicado: (2025)
por: Han, Woojung, et al.
Publicado: (2025)
Layer-Wise Modality Decomposition for Interpretable Multimodal Sensor Fusion
por: Park, Jaehyun, et al.
Publicado: (2025)
por: Park, Jaehyun, et al.
Publicado: (2025)
Human-Guided Shade Artifact Suppression in CBCT-to-MDCT Translation via Schrödinger Bridge with Conditional Diffusion
por: Kang, Sung Ho, et al.
Publicado: (2025)
por: Kang, Sung Ho, et al.
Publicado: (2025)
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
por: Xiao, Junhao, et al.
Publicado: (2026)
por: Xiao, Junhao, et al.
Publicado: (2026)
ODPG: Outfitting Diffusion with Pose Guided Condition
por: Lee, Seohyun, et al.
Publicado: (2025)
por: Lee, Seohyun, et al.
Publicado: (2025)
AH-OCDA: Amplitude-based Curriculum Learning and Hopfield Segmentation Model for Open Compound Domain Adaptation
por: Choi, Jaehyun, et al.
Publicado: (2024)
por: Choi, Jaehyun, et al.
Publicado: (2024)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
por: Burgess, James, et al.
Publicado: (2026)
por: Burgess, James, et al.
Publicado: (2026)
ART-VITON: Measurement-Guided Latent Diffusion for Artifact-Free Virtual Try-On
por: Park, Junseo, et al.
Publicado: (2025)
por: Park, Junseo, et al.
Publicado: (2025)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
por: Shi, Chufan, et al.
Publicado: (2026)
por: Shi, Chufan, et al.
Publicado: (2026)
Ejemplares similares
-
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
por: Kim, Minkyu, et al.
Publicado: (2026) -
Refining Visual Artifacts in Diffusion Models via Explainable AI-based Flaw Activation Maps
por: Lee, Seoyeon, et al.
Publicado: (2025) -
Test-time Alignment of Diffusion Models without Reward Over-optimization
por: Kim, Sunwoo, et al.
Publicado: (2025) -
Active Learning for Continual Learning: Keeping the Past Alive in the Present
por: Park, Jaehyun, et al.
Publicado: (2025) -
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
por: Park, Dongmin, et al.
Publicado: (2024)