Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: de Margerie, Anatole Jacquin, Roger, Alexis, Rish, Irina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917141113995264
author de Margerie, Anatole Jacquin
Roger, Alexis
Rish, Irina
author_facet de Margerie, Anatole Jacquin
Roger, Alexis
Rish, Irina
contents Reproducibility remains a cornerstone of scientific progress, yet complex multimodal models often lack transparent implementation details and accessible training infrastructure. In this work, we present a detailed reproduction and critical analysis of the Monkey Vision-Language Model (VLM) (Li et al. 2023b) published in CVPR24, a recent approach to high-resolution image understanding via image tiling. The original paper proposed splitting large images into tiles to recover fine-grained visual details while maintaining computational efficiency. Our study replicates this strategy using open checkpoints and reimplements the training pipeline. We confirm the key finding of the original Monkey VLM work, namely that tiling effectively recovers local details. We then extend this work further, by investigating the effect of the inclusion of the global context, which provide practical insights for future high-resolution multimodal modeling. However, we also report deviations in the results, with the magnitude of these effects depending heavily on task type and tile granularity.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11167
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
de Margerie, Anatole Jacquin
Roger, Alexis
Rish, Irina
Computer Vision and Pattern Recognition
Artificial Intelligence
Reproducibility remains a cornerstone of scientific progress, yet complex multimodal models often lack transparent implementation details and accessible training infrastructure. In this work, we present a detailed reproduction and critical analysis of the Monkey Vision-Language Model (VLM) (Li et al. 2023b) published in CVPR24, a recent approach to high-resolution image understanding via image tiling. The original paper proposed splitting large images into tiles to recover fine-grained visual details while maintaining computational efficiency. Our study replicates this strategy using open checkpoints and reimplements the training pipeline. We confirm the key finding of the original Monkey VLM work, namely that tiling effectively recovers local details. We then extend this work further, by investigating the effect of the inclusion of the global context, which provide practical insights for future high-resolution multimodal modeling. However, we also report deviations in the results, with the magnitude of these effects depending heavily on task type and tile granularity.
title Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.11167