Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Hanwen, Lin, Zexin, Deng, Yixuan, Ji, Xiaoqiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models
by: Lou, Ange, et al.
Published: (2026)
by: Lou, Ange, et al.
Published: (2026)
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
by: Saxena, Rohit, et al.
Published: (2026)
by: Saxena, Rohit, et al.
Published: (2026)
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
by: Jiang, Chaoya, et al.
Published: (2025)
by: Jiang, Chaoya, et al.
Published: (2025)
Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
by: Singhania, Aditi, et al.
Published: (2025)
by: Singhania, Aditi, et al.
Published: (2025)
Case-Aware Medical Image Classification with Multimodal Knowledge Graphs and Reliability-Guided Refinement
by: Xu, Yiming, et al.
Published: (2026)
by: Xu, Yiming, et al.
Published: (2026)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
by: Lin, Jiawen, et al.
Published: (2025)
by: Lin, Jiawen, et al.
Published: (2025)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
by: Xue, Qiyao, et al.
Published: (2024)
by: Xue, Qiyao, et al.
Published: (2024)
Dynamic VLM-Guided Negative Prompting for Diffusion Models
by: Chang, Hoyeon, et al.
Published: (2025)
by: Chang, Hoyeon, et al.
Published: (2025)
Towards Benchmarking and Evaluating Deepfake Detection
by: Lin, Chenhao, et al.
Published: (2022)
by: Lin, Chenhao, et al.
Published: (2022)
Trust but Verify: Programmatic VLM Evaluation in the Wild
by: Prabhu, Viraj, et al.
Published: (2024)
by: Prabhu, Viraj, et al.
Published: (2024)
AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
by: Liu, Guimeng, et al.
Published: (2025)
by: Liu, Guimeng, et al.
Published: (2025)
Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
by: Guo, Junrong, et al.
Published: (2026)
by: Guo, Junrong, et al.
Published: (2026)
IRNet: Iterative Refinement Network for Noisy Partial Label Learning
by: Lian, Zheng, et al.
Published: (2022)
by: Lian, Zheng, et al.
Published: (2022)
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
by: Xu, Chengzhi, et al.
Published: (2025)
by: Xu, Chengzhi, et al.
Published: (2025)
CLAMP: Contrastive Learning with Adaptive Multi-loss and Progressive Fusion for Multimodal Aspect-Based Sentiment Analysis
by: He, Xiaoqiang
Published: (2025)
by: He, Xiaoqiang
Published: (2025)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
by: Ji, Yicheng, et al.
Published: (2025)
by: Ji, Yicheng, et al.
Published: (2025)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
by: Jeong, Suchae, et al.
Published: (2025)
by: Jeong, Suchae, et al.
Published: (2025)
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
by: Chen, Ce, et al.
Published: (2026)
by: Chen, Ce, et al.
Published: (2026)
A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards
by: Patel, Shivansh, et al.
Published: (2025)
by: Patel, Shivansh, et al.
Published: (2025)
Unifying VLM-Guided Flow Matching and Spectral Anomaly Detection for Interpretable Veterinary Diagnosis
by: Wang, Pu, et al.
Published: (2026)
by: Wang, Pu, et al.
Published: (2026)
\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
by: Lei, Bin, et al.
Published: (2025)
by: Lei, Bin, et al.
Published: (2025)
Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference
by: Mao, Xiaowei, et al.
Published: (2026)
by: Mao, Xiaowei, et al.
Published: (2026)
CIARD: Cyclic Iterative Adversarial Robustness Distillation
by: Lu, Liming, et al.
Published: (2025)
by: Lu, Liming, et al.
Published: (2025)
BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
by: Wang, Shengao, et al.
Published: (2025)
by: Wang, Shengao, et al.
Published: (2025)
GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-based VLM Agent Training
by: Wei, Tong, et al.
Published: (2025)
by: Wei, Tong, et al.
Published: (2025)
SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
by: Jain, Chelsi, et al.
Published: (2025)
by: Jain, Chelsi, et al.
Published: (2025)
InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative Refinement
by: Zou, Yude, et al.
Published: (2026)
by: Zou, Yude, et al.
Published: (2026)
VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT
by: Kang, Hyeonsu, et al.
Published: (2025)
by: Kang, Hyeonsu, et al.
Published: (2025)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
by: Cho, Jaemin, et al.
Published: (2023)
by: Cho, Jaemin, et al.
Published: (2023)
ARIW-Framework: Adaptive Robust Iterative Watermarking Framework
by: Wu, Shaowu, et al.
Published: (2025)
by: Wu, Shaowu, et al.
Published: (2025)
How to Evaluate and Refine your CAM
by: Domeniconi, Luca, et al.
Published: (2026)
by: Domeniconi, Luca, et al.
Published: (2026)
Sim2Radar: Toward Bridging the Radar Sim-to-Real Gap with VLM-Guided Scene Reconstruction
by: Bejerano, Emily, et al.
Published: (2026)
by: Bejerano, Emily, et al.
Published: (2026)
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving
by: Huang, Zilin, et al.
Published: (2026)
by: Huang, Zilin, et al.
Published: (2026)
MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement
by: He, Xu, et al.
Published: (2024)
by: He, Xu, et al.
Published: (2024)
edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
by: Qian, Chen, et al.
Published: (2025)
by: Qian, Chen, et al.
Published: (2025)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
by: Singh, Aditya Kumar, et al.
Published: (2026)
by: Singh, Aditya Kumar, et al.
Published: (2026)
Perceptual Group Tokenizer: Building Perception with Iterative Grouping
by: Deng, Zhiwei, et al.
Published: (2023)
by: Deng, Zhiwei, et al.
Published: (2023)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
by: He, Chiyuan, et al.
Published: (2025)
by: He, Chiyuan, et al.
Published: (2025)
Similar Items
-
VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models
by: Lou, Ange, et al.
Published: (2026) -
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
by: Saxena, Rohit, et al.
Published: (2026) -
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
by: Jiang, Chaoya, et al.
Published: (2025) -
Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
by: Singhania, Aditi, et al.
Published: (2025) -
Case-Aware Medical Image Classification with Multimodal Knowledge Graphs and Reliability-Guided Refinement
by: Xu, Yiming, et al.
Published: (2026)