Saved in:
| Main Authors: | Yu, Xuan, Guan, Dayan, Gu, Yanfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.01663 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026)
by: Pei, Siqi, et al.
Published: (2026)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
by: Erzurumlu, Yunus Talha, et al.
Published: (2026)
by: Erzurumlu, Yunus Talha, et al.
Published: (2026)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
by: Luan, Bozhi, et al.
Published: (2024)
by: Luan, Bozhi, et al.
Published: (2024)
AttZoom: Attention Zoom for Better Visual Features
by: DeAlcala, Daniel, et al.
Published: (2025)
by: DeAlcala, Daniel, et al.
Published: (2025)
ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
by: Shen, Haozhan, et al.
Published: (2024)
by: Shen, Haozhan, et al.
Published: (2024)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
by: Wei, Lai, et al.
Published: (2026)
by: Wei, Lai, et al.
Published: (2026)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
by: Liu, Xianjie, et al.
Published: (2025)
by: Liu, Xianjie, et al.
Published: (2025)
Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
by: Yang, Jiashu, et al.
Published: (2025)
by: Yang, Jiashu, et al.
Published: (2025)
ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
by: Liu, Ruixun, et al.
Published: (2025)
by: Liu, Ruixun, et al.
Published: (2025)
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
by: Jiang, Zhiyuan, et al.
Published: (2025)
by: Jiang, Zhiyuan, et al.
Published: (2025)
GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance
by: Shi, Jiale, et al.
Published: (2026)
by: Shi, Jiale, et al.
Published: (2026)
Self-Supervised Learning for Real-World Super-Resolution from Dual and Multiple Zoomed Observations
by: Zhang, Zhilu, et al.
Published: (2024)
by: Zhang, Zhilu, et al.
Published: (2024)
Dual-Camera Smooth Zoom on Mobile Phones
by: Wu, Renlong, et al.
Published: (2024)
by: Wu, Renlong, et al.
Published: (2024)
PTZ-Calib: Robust Pan-Tilt-Zoom Camera Calibration
by: Guo, Jinhui, et al.
Published: (2025)
by: Guo, Jinhui, et al.
Published: (2025)
Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding
by: Li, Chenglin, et al.
Published: (2025)
by: Li, Chenglin, et al.
Published: (2025)
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
by: Thapa, Rahul, et al.
Published: (2024)
by: Thapa, Rahul, et al.
Published: (2024)
Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
by: Guo, Jiayi, et al.
Published: (2026)
by: Guo, Jiayi, et al.
Published: (2026)
Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment
by: Kim, Bryan Sangwoo, et al.
Published: (2025)
by: Kim, Bryan Sangwoo, et al.
Published: (2025)
Temporal Zoom Networks: Distance Regression and Continuous Depth for Efficient Action Localization
by: Shihab, Ibne Farabi, et al.
Published: (2025)
by: Shihab, Ibne Farabi, et al.
Published: (2025)
Zoom and Shift are All You Need
by: Qin, Jiahao
Published: (2024)
by: Qin, Jiahao
Published: (2024)
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
by: Zhou, Dewei, et al.
Published: (2026)
by: Zhou, Dewei, et al.
Published: (2026)
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
Seeing the Unseen: Zooming in the Dark with Event Cameras
by: Kai, Dachun, et al.
Published: (2026)
by: Kai, Dachun, et al.
Published: (2026)
SpectralZoom: Efficient Segmentation with an Adaptive Hyperspectral Camera
by: Arnold, Jackson, et al.
Published: (2024)
by: Arnold, Jackson, et al.
Published: (2024)
MedReason-R1: Learning to Reason for CT Diagnosis with Reinforcement Learning and Local Zoom
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
UltraZoom: Generating Gigapixel Images from Regular Photos
by: Ma, Jingwei, et al.
Published: (2025)
by: Ma, Jingwei, et al.
Published: (2025)
Efficient Hybrid Zoom using Camera Fusion on Mobile Phones
by: Wu, Xiaotong, et al.
Published: (2024)
by: Wu, Xiaotong, et al.
Published: (2024)
ATAC-Net: Zoomed view works better for Anomaly Detection
by: Gupta, Shaurya, et al.
Published: (2024)
by: Gupta, Shaurya, et al.
Published: (2024)
Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
by: Moser, Brian B., et al.
Published: (2024)
by: Moser, Brian B., et al.
Published: (2024)
Rethinking the Evaluation of Visible and Infrared Image Fusion
by: Guan, Dayan, et al.
Published: (2024)
by: Guan, Dayan, et al.
Published: (2024)
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
by: Ma, Yunchuan, et al.
Published: (2026)
by: Ma, Yunchuan, et al.
Published: (2026)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)
by: Zhang, Xintong, et al.
Published: (2025)
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models
by: Shi, Yuheng, et al.
Published: (2026)
by: Shi, Yuheng, et al.
Published: (2026)
WonderZoom: Multi-Scale 3D World Generation
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification Approach
by: Cai, Lvpan, et al.
Published: (2025)
by: Cai, Lvpan, et al.
Published: (2025)
Adaptive Image Zoom-in with Bounding Box Transformation for UAV Object Detection
by: Wang, Tao, et al.
Published: (2026)
by: Wang, Tao, et al.
Published: (2026)
Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning
by: Liang, Guoqiang, et al.
Published: (2026)
by: Liang, Guoqiang, et al.
Published: (2026)
ZoomLDM: Latent Diffusion Model for multi-scale image generation
by: Yellapragada, Srikar, et al.
Published: (2024)
by: Yellapragada, Srikar, et al.
Published: (2024)
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
by: Zhou, Yue, et al.
Published: (2026)
by: Zhou, Yue, et al.
Published: (2026)
Similar Items
-
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026) -
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025) -
Just Zoom In: Cross-View Geo-Localization via Autoregressive Zooming
by: Erzurumlu, Yunus Talha, et al.
Published: (2026) -
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
by: Luan, Bozhi, et al.
Published: (2024) -
AttZoom: Attention Zoom for Better Visual Features
by: DeAlcala, Daniel, et al.
Published: (2025)