Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Thapa, Rahul, Chen, Kezhen, Covert, Ian, Chalamala, Rahul, Athiwaratkun, Ben, Song, Shuaiwen Leon, Zou, James |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SMIR: Efficient Synthetic Data Pipeline To Improve Multi-Image Reasoning
by: Li, Andrew, et al.
Published: (2025)
by: Li, Andrew, et al.
Published: (2025)
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
by: Thapa, Rahul, et al.
Published: (2025)
by: Thapa, Rahul, et al.
Published: (2025)
Locality Alignment Improves Vision-Language Models
by: Covert, Ian, et al.
Published: (2024)
by: Covert, Ian, et al.
Published: (2024)
ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
by: Liu, Ruixun, et al.
Published: (2025)
by: Liu, Ruixun, et al.
Published: (2025)
MURE: Hierarchical Multi-Resolution Encoding via Vision-Language Models for Visual Document Retrieval
by: Zhu, Fengbin, et al.
Published: (2026)
by: Zhu, Fengbin, et al.
Published: (2026)
A PCA based Keypoint Tracking Approach to Automated Facial Expressions Encoding
by: Tripathi, Shivansh Chandra, et al.
Published: (2024)
by: Tripathi, Shivansh Chandra, et al.
Published: (2024)
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
by: Yu, Xuan, et al.
Published: (2025)
by: Yu, Xuan, et al.
Published: (2025)
Revisiting Multimodal Positional Encoding in Vision-Language Models
by: Huang, Jie, et al.
Published: (2025)
by: Huang, Jie, et al.
Published: (2025)
CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
by: Anand, Neeraj, et al.
Published: (2026)
by: Anand, Neeraj, et al.
Published: (2026)
MicroVision: An Open Dataset and Benchmark Models for Detecting Vulnerable Road Users and Micromobility Vehicles
by: Rasch, Alexander, et al.
Published: (2026)
by: Rasch, Alexander, et al.
Published: (2026)
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
by: Unmesh, Asim, et al.
Published: (2026)
by: Unmesh, Asim, et al.
Published: (2026)
HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
by: Liu, Xianjie, et al.
Published: (2025)
by: Liu, Xianjie, et al.
Published: (2025)
A Review of 3D Object Detection with Vision-Language Models
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Explanation-Aware Learning for Enhanced Interpretability in Biomedical Imaging
by: Faruqui, Zubair, et al.
Published: (2026)
by: Faruqui, Zubair, et al.
Published: (2026)
EventZoom: A Progressive Approach to Event-Based Data Augmentation for Enhanced Neuromorphic Vision
by: Dong, Yiting, et al.
Published: (2024)
by: Dong, Yiting, et al.
Published: (2024)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
by: Lu, Jiaying, et al.
Published: (2023)
by: Lu, Jiaying, et al.
Published: (2023)
Interpreting Neurons in Deep Vision Networks with Language Models
by: Bai, Nicholas, et al.
Published: (2024)
by: Bai, Nicholas, et al.
Published: (2024)
Temporal-Anchor3DLane: Enhanced 3D Lane Detection with Multi-Task Losses and LSTM Fusion
by: Suhas, D. Shainu, et al.
Published: (2025)
by: Suhas, D. Shainu, et al.
Published: (2025)
ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
by: Shen, Haozhan, et al.
Published: (2024)
by: Shen, Haozhan, et al.
Published: (2024)
OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
by: Lu, Pan, et al.
Published: (2025)
by: Lu, Pan, et al.
Published: (2025)
Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models
by: Zhang, Yunkai, et al.
Published: (2026)
by: Zhang, Yunkai, et al.
Published: (2026)
Computer Vision based Automated Quantification of Agricultural Sprayers Boom Displacement
by: Dalal, Aryan Singh, et al.
Published: (2025)
by: Dalal, Aryan Singh, et al.
Published: (2025)
Characterizing and Optimizing the Spatial Kernel of Multi Resolution Hash Encodings
by: Dai, Tianxiang, et al.
Published: (2026)
by: Dai, Tianxiang, et al.
Published: (2026)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2025)
by: Tian, Xinyu, et al.
Published: (2025)
MZEN: Multi-Zoom Enhanced NeRF for 3-D Reconstruction with Unknown Camera Poses
by: Park, Jong-Ik, et al.
Published: (2025)
by: Park, Jong-Ik, et al.
Published: (2025)
Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
by: Chen, Yule, et al.
Published: (2025)
by: Chen, Yule, et al.
Published: (2025)
Test-time Alignment-Enhanced Adapter for Vision-Language Models
by: Tong, Baoshun, et al.
Published: (2024)
by: Tong, Baoshun, et al.
Published: (2024)
Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
by: KV, Karthikeya
Published: (2025)
by: KV, Karthikeya
Published: (2025)
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models
by: Zhou, Xirui, et al.
Published: (2025)
by: Zhou, Xirui, et al.
Published: (2025)
On Evaluation of Vision Datasets and Models using Human Competency Frameworks
by: Ramachandran, Rahul, et al.
Published: (2024)
by: Ramachandran, Rahul, et al.
Published: (2024)
ClipGrader: Leveraging Vision-Language Models for Robust Label Quality Assessment in Object Detection
by: Lu, Hong, et al.
Published: (2025)
by: Lu, Hong, et al.
Published: (2025)
Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models
by: Imran, Muhammad, et al.
Published: (2025)
by: Imran, Muhammad, et al.
Published: (2025)
FastVLM: Efficient Vision Encoding for Vision Language Models
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
UV-Mamba: A DCN-Enhanced State Space Model for Urban Village Boundary Identification in High-Resolution Remote Sensing Images
by: Li, Lulin, et al.
Published: (2024)
by: Li, Lulin, et al.
Published: (2024)
OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial Optimization
by: Meng, Xiang, et al.
Published: (2024)
by: Meng, Xiang, et al.
Published: (2024)
SeG-SR: Integrating Semantic Knowledge into Remote Sensing Image Super-Resolution via Vision-Language Model
by: Chen, Bowen, et al.
Published: (2025)
by: Chen, Bowen, et al.
Published: (2025)
Self-Supervised Learning for Real-World Super-Resolution from Dual and Multiple Zoomed Observations
by: Zhang, Zhilu, et al.
Published: (2024)
by: Zhang, Zhilu, et al.
Published: (2024)
Human-inspired Global-to-Parallel Multi-scale Encoding for Lightweight Vision Models
by: Xu, Wei
Published: (2026)
by: Xu, Wei
Published: (2026)
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
by: Yuan, Kun, et al.
Published: (2025)
by: Yuan, Kun, et al.
Published: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
Similar Items
-
SMIR: Efficient Synthetic Data Pipeline To Improve Multi-Image Reasoning
by: Li, Andrew, et al.
Published: (2025) -
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
by: Thapa, Rahul, et al.
Published: (2025) -
Locality Alignment Improves Vision-Language Models
by: Covert, Ian, et al.
Published: (2024) -
ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
by: Liu, Ruixun, et al.
Published: (2025) -
MURE: Hierarchical Multi-Resolution Encoding via Vision-Language Models for Visual Document Retrieval
by: Zhu, Fengbin, et al.
Published: (2026)