Zoom Consistency: A Free Confidence Signal in Multi-Step Visual Grounding Pipelines
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Keon, Chelikavada, Krish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AttZoom: Attention Zoom for Better Visual Features
von: DeAlcala, Daniel, et al.
Veröffentlicht: (2025)
von: DeAlcala, Daniel, et al.
Veröffentlicht: (2025)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
von: Pei, Siqi, et al.
Veröffentlicht: (2026)
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2025)
Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
von: Dai, Ming, et al.
Veröffentlicht: (2025)
von: Dai, Ming, et al.
Veröffentlicht: (2025)
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
von: Zhou, Yue, et al.
Veröffentlicht: (2026)
von: Zhou, Yue, et al.
Veröffentlicht: (2026)
Training-Free Consistency Pipeline for Fashion Repose
von: Aghilar, Potito, et al.
Veröffentlicht: (2025)
von: Aghilar, Potito, et al.
Veröffentlicht: (2025)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
von: Tang, Fei, et al.
Veröffentlicht: (2026)
von: Tang, Fei, et al.
Veröffentlicht: (2026)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
von: Li, Zejun, et al.
Veröffentlicht: (2024)
von: Li, Zejun, et al.
Veröffentlicht: (2024)
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
SFUOD: Source-Free Unknown Object Detection
von: Park, Keon-Hee, et al.
Veröffentlicht: (2025)
von: Park, Keon-Hee, et al.
Veröffentlicht: (2025)
CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
von: Zhang, Ruoxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruoxuan, et al.
Veröffentlicht: (2025)
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
von: Thapa, Rahul, et al.
Veröffentlicht: (2024)
von: Thapa, Rahul, et al.
Veröffentlicht: (2024)
STEVE: A Step Verification Pipeline for Computer-use Agent Training
von: Lu, Fanbin, et al.
Veröffentlicht: (2025)
von: Lu, Fanbin, et al.
Veröffentlicht: (2025)
Zoom and Shift are All You Need
von: Qin, Jiahao
Veröffentlicht: (2024)
von: Qin, Jiahao
Veröffentlicht: (2024)
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
von: Wang, Jingchao, et al.
Veröffentlicht: (2025)
MEET: A Million-Scale Dataset for Fine-Grained Geospatial Scene Classification with Zoom-Free Remote Sensing Imagery
von: Li, Yansheng, et al.
Veröffentlicht: (2025)
von: Li, Yansheng, et al.
Veröffentlicht: (2025)
WonderZoom: Multi-Scale 3D World Generation
von: Cao, Jin, et al.
Veröffentlicht: (2025)
von: Cao, Jin, et al.
Veröffentlicht: (2025)
A Simple and Effective Temporal Grounding Pipeline for Basketball Broadcast Footage
von: Harris, Levi
Veröffentlicht: (2024)
von: Harris, Levi
Veröffentlicht: (2024)
LCV2: An Efficient Pretraining-Free Framework for Grounded Visual Question Answering
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
von: Chen, Yuhan, et al.
Veröffentlicht: (2024)
Seeing the Unseen: Zooming in the Dark with Event Cameras
von: Kai, Dachun, et al.
Veröffentlicht: (2026)
von: Kai, Dachun, et al.
Veröffentlicht: (2026)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
AnatomicalNets: A Multi-Structure Segmentation and Contour-Based Distance Estimation Pipeline for Clinically Grounded Lung Cancer T-Staging
von: Chowdhury, Saniah Kayenat, et al.
Veröffentlicht: (2025)
von: Chowdhury, Saniah Kayenat, et al.
Veröffentlicht: (2025)
Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
von: Jain, Chayan, et al.
Veröffentlicht: (2025)
von: Jain, Chayan, et al.
Veröffentlicht: (2025)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
von: Wu, Qianhui, et al.
Veröffentlicht: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
von: Lim, Byeonggeuk, et al.
Veröffentlicht: (2026)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
GreenEye: Development of Real-Time Traffic Signal Recognition System for Visual Impairments
von: Kim, Danu
Veröffentlicht: (2024)
von: Kim, Danu
Veröffentlicht: (2024)
Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement
von: Xu, Jian, et al.
Veröffentlicht: (2025)
von: Xu, Jian, et al.
Veröffentlicht: (2025)
AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
von: Li, Haocheng, et al.
Veröffentlicht: (2026)
von: Li, Haocheng, et al.
Veröffentlicht: (2026)
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
von: Dai, Ming, et al.
Veröffentlicht: (2024)
von: Dai, Ming, et al.
Veröffentlicht: (2024)
Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
YOLO-Based Pipeline Monitoring in Challenging Visual Environments
von: Dhungana, Pragya, et al.
Veröffentlicht: (2025)
von: Dhungana, Pragya, et al.
Veröffentlicht: (2025)
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
von: Lv, Guannan, et al.
Veröffentlicht: (2026)
von: Lv, Guannan, et al.
Veröffentlicht: (2026)
PathGLS: Evaluating Pathology Vision-Language Models without Ground Truth through Multi-Dimensional Consistency
von: Chen, Minbing, et al.
Veröffentlicht: (2026)
von: Chen, Minbing, et al.
Veröffentlicht: (2026)
Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment
von: Kim, Bryan Sangwoo, et al.
Veröffentlicht: (2025)
von: Kim, Bryan Sangwoo, et al.
Veröffentlicht: (2025)
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
von: Yao, Zhengjian, et al.
Veröffentlicht: (2026)
von: Yao, Zhengjian, et al.
Veröffentlicht: (2026)
A Proxy Consistency Loss for Grounded Fusion of Earth Observation and Location Encoders
von: Wang, Zhongying, et al.
Veröffentlicht: (2026)
von: Wang, Zhongying, et al.
Veröffentlicht: (2026)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
Visual Position Prompt for MLLM based Visual Grounding
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
Towards Understanding Visual Grounding in Visual Language Models
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2025)
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AttZoom: Attention Zoom for Better Visual Features
von: DeAlcala, Daniel, et al.
Veröffentlicht: (2025) -
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
von: Pei, Siqi, et al.
Veröffentlicht: (2026) -
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2025) -
Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
von: Dai, Ming, et al.
Veröffentlicht: (2025) -
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
von: Zhou, Yue, et al.
Veröffentlicht: (2026)