Salvato in:
| Autori principali: | Kim, Keon, Chelikavada, Krish |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2604.15376 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AttZoom: Attention Zoom for Better Visual Features
di: DeAlcala, Daniel, et al.
Pubblicazione: (2025)
di: DeAlcala, Daniel, et al.
Pubblicazione: (2025)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
di: Pei, Siqi, et al.
Pubblicazione: (2026)
di: Pei, Siqi, et al.
Pubblicazione: (2026)
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
di: Jiang, Zhiyuan, et al.
Pubblicazione: (2025)
di: Jiang, Zhiyuan, et al.
Pubblicazione: (2025)
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
di: Zhou, Yue, et al.
Pubblicazione: (2026)
di: Zhou, Yue, et al.
Pubblicazione: (2026)
Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
di: Dai, Ming, et al.
Pubblicazione: (2025)
di: Dai, Ming, et al.
Pubblicazione: (2025)
Training-Free Consistency Pipeline for Fashion Repose
di: Aghilar, Potito, et al.
Pubblicazione: (2025)
di: Aghilar, Potito, et al.
Pubblicazione: (2025)
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding
di: Tang, Fei, et al.
Pubblicazione: (2026)
di: Tang, Fei, et al.
Pubblicazione: (2026)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
di: Li, Zejun, et al.
Pubblicazione: (2024)
di: Li, Zejun, et al.
Pubblicazione: (2024)
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
di: Erzurumlu, Yunus Talha, et al.
Pubblicazione: (2026)
di: Erzurumlu, Yunus Talha, et al.
Pubblicazione: (2026)
SFUOD: Source-Free Unknown Object Detection
di: Park, Keon-Hee, et al.
Pubblicazione: (2025)
di: Park, Keon-Hee, et al.
Pubblicazione: (2025)
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
di: Thapa, Rahul, et al.
Pubblicazione: (2024)
di: Thapa, Rahul, et al.
Pubblicazione: (2024)
CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
Zoom and Shift are All You Need
di: Qin, Jiahao
Pubblicazione: (2024)
di: Qin, Jiahao
Pubblicazione: (2024)
STEVE: A Step Verification Pipeline for Computer-use Agent Training
di: Lu, Fanbin, et al.
Pubblicazione: (2025)
di: Lu, Fanbin, et al.
Pubblicazione: (2025)
WonderZoom: Multi-Scale 3D World Generation
di: Cao, Jin, et al.
Pubblicazione: (2025)
di: Cao, Jin, et al.
Pubblicazione: (2025)
MEET: A Million-Scale Dataset for Fine-Grained Geospatial Scene Classification with Zoom-Free Remote Sensing Imagery
di: Li, Yansheng, et al.
Pubblicazione: (2025)
di: Li, Yansheng, et al.
Pubblicazione: (2025)
Seeing the Unseen: Zooming in the Dark with Event Cameras
di: Kai, Dachun, et al.
Pubblicazione: (2026)
di: Kai, Dachun, et al.
Pubblicazione: (2026)
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
di: Wang, Jingchao, et al.
Pubblicazione: (2025)
di: Wang, Jingchao, et al.
Pubblicazione: (2025)
A Simple and Effective Temporal Grounding Pipeline for Basketball Broadcast Footage
di: Harris, Levi
Pubblicazione: (2024)
di: Harris, Levi
Pubblicazione: (2024)
LCV2: An Efficient Pretraining-Free Framework for Grounded Visual Question Answering
di: Chen, Yuhan, et al.
Pubblicazione: (2024)
di: Chen, Yuhan, et al.
Pubblicazione: (2024)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
di: Rahman, Md Ashikur, et al.
Pubblicazione: (2026)
di: Rahman, Md Ashikur, et al.
Pubblicazione: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
di: Wu, Qianhui, et al.
Pubblicazione: (2025)
di: Wu, Qianhui, et al.
Pubblicazione: (2025)
Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
di: Jain, Chayan, et al.
Pubblicazione: (2025)
di: Jain, Chayan, et al.
Pubblicazione: (2025)
Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment
di: Kim, Bryan Sangwoo, et al.
Pubblicazione: (2025)
di: Kim, Bryan Sangwoo, et al.
Pubblicazione: (2025)
AnatomicalNets: A Multi-Structure Segmentation and Contour-Based Distance Estimation Pipeline for Clinically Grounded Lung Cancer T-Staging
di: Chowdhury, Saniah Kayenat, et al.
Pubblicazione: (2025)
di: Chowdhury, Saniah Kayenat, et al.
Pubblicazione: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
di: Lim, Byeonggeuk, et al.
Pubblicazione: (2026)
di: Lim, Byeonggeuk, et al.
Pubblicazione: (2026)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
di: Kang, Seil, et al.
Pubblicazione: (2025)
di: Kang, Seil, et al.
Pubblicazione: (2025)
Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding
di: Li, Chenglin, et al.
Pubblicazione: (2025)
di: Li, Chenglin, et al.
Pubblicazione: (2025)
Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement
di: Xu, Jian, et al.
Pubblicazione: (2025)
di: Xu, Jian, et al.
Pubblicazione: (2025)
GreenEye: Development of Real-Time Traffic Signal Recognition System for Visual Impairments
di: Kim, Danu
Pubblicazione: (2024)
di: Kim, Danu
Pubblicazione: (2024)
AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
di: Li, Haocheng, et al.
Pubblicazione: (2026)
di: Li, Haocheng, et al.
Pubblicazione: (2026)
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
di: Dai, Ming, et al.
Pubblicazione: (2024)
di: Dai, Ming, et al.
Pubblicazione: (2024)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
di: Wei, Lai, et al.
Pubblicazione: (2026)
di: Wei, Lai, et al.
Pubblicazione: (2026)
YOLO-Based Pipeline Monitoring in Challenging Visual Environments
di: Dhungana, Pragya, et al.
Pubblicazione: (2025)
di: Dhungana, Pragya, et al.
Pubblicazione: (2025)
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
di: Lv, Guannan, et al.
Pubblicazione: (2026)
di: Lv, Guannan, et al.
Pubblicazione: (2026)
PathGLS: Evaluating Pathology Vision-Language Models without Ground Truth through Multi-Dimensional Consistency
di: Chen, Minbing, et al.
Pubblicazione: (2026)
di: Chen, Minbing, et al.
Pubblicazione: (2026)
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
di: Yao, Zhengjian, et al.
Pubblicazione: (2026)
di: Yao, Zhengjian, et al.
Pubblicazione: (2026)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
di: Kim, Keuntae, et al.
Pubblicazione: (2026)
di: Kim, Keuntae, et al.
Pubblicazione: (2026)
A Proxy Consistency Loss for Grounded Fusion of Earth Observation and Location Encoders
di: Wang, Zhongying, et al.
Pubblicazione: (2026)
di: Wang, Zhongying, et al.
Pubblicazione: (2026)
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
di: Sun, Yujing, et al.
Pubblicazione: (2025)
di: Sun, Yujing, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AttZoom: Attention Zoom for Better Visual Features
di: DeAlcala, Daniel, et al.
Pubblicazione: (2025) -
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
di: Pei, Siqi, et al.
Pubblicazione: (2026) -
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
di: Jiang, Zhiyuan, et al.
Pubblicazione: (2025) -
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
di: Zhou, Yue, et al.
Pubblicazione: (2026) -
Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
di: Dai, Ming, et al.
Pubblicazione: (2025)