ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Haozhan, Zhao, Kangjia, Zhao, Tiancheng, Xu, Ruochen, Zhang, Zilun, Zhu, Mingwei, Yin, Jianwei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
by: Shen, Haozhan, et al.
Published: (2026)
by: Shen, Haozhan, et al.
Published: (2026)
The Self-Improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities without External Scaffolding?
by: Sun, Yutao, et al.
Published: (2025)
by: Sun, Yutao, et al.
Published: (2025)
Talking to Yourself: Defying Forgetting in Large Language Models
by: Sun, Yutao, et al.
Published: (2026)
by: Sun, Yutao, et al.
Published: (2026)
ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG
by: Zhang, Zilun, et al.
Published: (2024)
by: Zhang, Zilun, et al.
Published: (2024)
GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
by: Zhao, Kangjia, et al.
Published: (2024)
by: Zhao, Kangjia, et al.
Published: (2024)
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
by: Shen, Haozhan, et al.
Published: (2025)
by: Shen, Haozhan, et al.
Published: (2025)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
by: Wei, Lai, et al.
Published: (2026)
by: Wei, Lai, et al.
Published: (2026)
RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
by: Zhang, Zilun, et al.
Published: (2023)
by: Zhang, Zilun, et al.
Published: (2023)
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing
by: Zhang, Zilun, et al.
Published: (2025)
by: Zhang, Zilun, et al.
Published: (2025)
Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression
by: Zhang, Zilun, et al.
Published: (2024)
by: Zhang, Zilun, et al.
Published: (2024)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
by: Zhang, Zilun, et al.
Published: (2022)
by: Zhang, Zilun, et al.
Published: (2022)
SRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation
by: Guo, Yulong, et al.
Published: (2025)
by: Guo, Yulong, et al.
Published: (2025)
ZoomTable: Interactive Exploration of Data Facts in Hierarchical Tables via Semantic Zooming
by: Chen, Qiyang, et al.
Published: (2026)
by: Chen, Qiyang, et al.
Published: (2026)
Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding
by: Li, Chenglin, et al.
Published: (2025)
by: Li, Chenglin, et al.
Published: (2025)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
by: Luan, Bozhi, et al.
Published: (2024)
by: Luan, Bozhi, et al.
Published: (2024)
Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning
by: Zhang, Zilun, et al.
Published: (2025)
by: Zhang, Zilun, et al.
Published: (2025)
AttZoom: Attention Zoom for Better Visual Features
by: DeAlcala, Daniel, et al.
Published: (2025)
by: DeAlcala, Daniel, et al.
Published: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
by: Yu, Xuan, et al.
Published: (2025)
by: Yu, Xuan, et al.
Published: (2025)
EventZoom: A Progressive Approach to Event-Based Data Augmentation for Enhanced Neuromorphic Vision
by: Dong, Yiting, et al.
Published: (2024)
by: Dong, Yiting, et al.
Published: (2024)
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
by: Jiang, Zhiyuan, et al.
Published: (2025)
by: Jiang, Zhiyuan, et al.
Published: (2025)
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
by: Erzurumlu, Yunus Talha, et al.
Published: (2026)
by: Erzurumlu, Yunus Talha, et al.
Published: (2026)
FIZZ: Factual Inconsistency Detection by Zoom-in Summary and Zoom-out Document
by: Yang, Joonho, et al.
Published: (2024)
by: Yang, Joonho, et al.
Published: (2024)
HeadZoom: Hands-Free Zooming and Panning for 2D Image Navigation Using Head Motion
by: Zhang, Kaining, et al.
Published: (2025)
by: Zhang, Kaining, et al.
Published: (2025)
Zoom Conference Dataset
by: Gudaparthi, Hemanth
Published: (2025)
by: Gudaparthi, Hemanth
Published: (2025)
Zooming in on discrete space
by: Vanzella, Daniel A. Turolla
Published: (2024)
by: Vanzella, Daniel A. Turolla
Published: (2024)
Code Semantic Zooming
by: Ba, Jinsheng, et al.
Published: (2025)
by: Ba, Jinsheng, et al.
Published: (2025)
Who's Watching You Zoom? Investigating Privacy of Third-Party Zoom Apps
by: Goenka, Saharsh, et al.
Published: (2025)
by: Goenka, Saharsh, et al.
Published: (2025)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026)
by: Pei, Siqi, et al.
Published: (2026)
To Zoom or not to Zoom: Assessment of synchronous online modality preferences and performance in an introductory undergraduate course
by: Rachel L. Veenstra Cott
Published: (2025)
by: Rachel L. Veenstra Cott
Published: (2025)
Adaptive Image Zoom-in with Bounding Box Transformation for UAV Object Detection
by: Wang, Tao, et al.
Published: (2026)
by: Wang, Tao, et al.
Published: (2026)
GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance
by: Shi, Jiale, et al.
Published: (2026)
by: Shi, Jiale, et al.
Published: (2026)
MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
by: Shen, Haozhan, et al.
Published: (2026)
by: Shen, Haozhan, et al.
Published: (2026)
Physics‐Inspired Diffractive Neural Networks for Zooming Image Edge Detection
by: Weijie Chang, et al.
Published: (2025)
by: Weijie Chang, et al.
Published: (2025)
Seeing the Unseen: Zooming in the Dark with Event Cameras
by: Kai, Dachun, et al.
Published: (2026)
by: Kai, Dachun, et al.
Published: (2026)
Zoom and Shift are All You Need
by: Qin, Jiahao
Published: (2024)
by: Qin, Jiahao
Published: (2024)
Equilibrium States for Random Zooming Systems
by: Bilbao, Rafael A., et al.
Published: (2023)
by: Bilbao, Rafael A., et al.
Published: (2023)
Equilibrium Stability for Open Zooming Systems
by: Bilbao, Rafael A., et al.
Published: (2025)
by: Bilbao, Rafael A., et al.
Published: (2025)
Equilibrium States for Open Zooming Systems
by: Santana, Eduardo
Published: (2020)
by: Santana, Eduardo
Published: (2020)
Zoom into Pre-School Story Hour.
by: Glaser, Ann, et al.
Published: (1982)
by: Glaser, Ann, et al.
Published: (1982)
Similar Items
-
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
by: Shen, Haozhan, et al.
Published: (2026) -
The Self-Improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities without External Scaffolding?
by: Sun, Yutao, et al.
Published: (2025) -
Talking to Yourself: Defying Forgetting in Large Language Models
by: Sun, Yutao, et al.
Published: (2026) -
ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG
by: Zhang, Zilun, et al.
Published: (2024) -
GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
by: Zhao, Kangjia, et al.
Published: (2024)