Thyme: Think Beyond Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yi-Fan, Lu, Xingyu, Yin, Shukang, Fu, Chaoyou, Chen, Wei, Hu, Xiao, Wen, Bin, Jiang, Kaiyu, Liu, Changyi, Zhang, Tianke, Fan, Haonan, Chen, Kaibing, Chen, Jiankang, Ding, Haojie, Tang, Kaiyu, Zhang, Zhang, Wang, Liang, Yang, Fan, Gao, Tingting, Zhou, Guorui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912539361673216
author Zhang, Yi-Fan
Lu, Xingyu
Yin, Shukang
Fu, Chaoyou
Chen, Wei
Hu, Xiao
Wen, Bin
Jiang, Kaiyu
Liu, Changyi
Zhang, Tianke
Fan, Haonan
Chen, Kaibing
Chen, Jiankang
Ding, Haojie
Tang, Kaiyu
Zhang, Zhang
Wang, Liang
Yang, Fan
Gao, Tingting
Zhou, Guorui
author_facet Zhang, Yi-Fan
Lu, Xingyu
Yin, Shukang
Fu, Chaoyou
Chen, Wei
Hu, Xiao
Wen, Bin
Jiang, Kaiyu
Liu, Changyi
Zhang, Tianke
Fan, Haonan
Chen, Kaibing
Chen, Jiankang
Ding, Haojie
Tang, Kaiyu
Zhang, Zhang
Wang, Liang
Yang, Fan
Gao, Tingting
Zhou, Guorui
contents Following OpenAI's introduction of the ``thinking with images'' concept, recent efforts have explored stimulating the use of visual information in the reasoning process to enhance model performance in perception and reasoning tasks. However, to the best of our knowledge, no open-source work currently offers a feature set as rich as proprietary models (O3), which can perform diverse image manipulations and simultaneously enhance logical reasoning capabilities through code. In this paper, we make a preliminary attempt in this direction by introducing Thyme (Think Beyond Images), a novel paradigm for enabling MLLMs to transcend existing ``think with images'' approaches by autonomously generating and executing diverse image processing and computational operations via executable code. This approach not only facilitates a rich, on-the-fly set of image manipulations (e.g., cropping, rotation, contrast enhancement) but also allows for mathematical computations, all while maintaining high autonomy in deciding when and how to apply these operations. We activate this capability through a two-stage training strategy: an initial SFT on a curated dataset of 500K samples to teach code generation, followed by a RL phase to refine decision-making. For the RL stage, we manually collect and design high-resolution question-answer pairs to increase the learning difficulty, and we propose GRPO-ATS (Group Relative Policy Optimization with Adaptive Temperature Sampling), an algorithm that applies distinct temperatures to text and code generation to balance reasoning exploration with code execution precision. We conduct extensive experimental analysis and ablation studies. Comprehensive evaluations on nearly 20 benchmarks show that Thyme yields significant and consistent performance gains, particularly in challenging high-resolution perception and complex reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11630
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thyme: Think Beyond Images
Zhang, Yi-Fan
Lu, Xingyu
Yin, Shukang
Fu, Chaoyou
Chen, Wei
Hu, Xiao
Wen, Bin
Jiang, Kaiyu
Liu, Changyi
Zhang, Tianke
Fan, Haonan
Chen, Kaibing
Chen, Jiankang
Ding, Haojie
Tang, Kaiyu
Zhang, Zhang
Wang, Liang
Yang, Fan
Gao, Tingting
Zhou, Guorui
Computer Vision and Pattern Recognition
Following OpenAI's introduction of the ``thinking with images'' concept, recent efforts have explored stimulating the use of visual information in the reasoning process to enhance model performance in perception and reasoning tasks. However, to the best of our knowledge, no open-source work currently offers a feature set as rich as proprietary models (O3), which can perform diverse image manipulations and simultaneously enhance logical reasoning capabilities through code. In this paper, we make a preliminary attempt in this direction by introducing Thyme (Think Beyond Images), a novel paradigm for enabling MLLMs to transcend existing ``think with images'' approaches by autonomously generating and executing diverse image processing and computational operations via executable code. This approach not only facilitates a rich, on-the-fly set of image manipulations (e.g., cropping, rotation, contrast enhancement) but also allows for mathematical computations, all while maintaining high autonomy in deciding when and how to apply these operations. We activate this capability through a two-stage training strategy: an initial SFT on a curated dataset of 500K samples to teach code generation, followed by a RL phase to refine decision-making. For the RL stage, we manually collect and design high-resolution question-answer pairs to increase the learning difficulty, and we propose GRPO-ATS (Group Relative Policy Optimization with Adaptive Temperature Sampling), an algorithm that applies distinct temperatures to text and code generation to balance reasoning exploration with code execution precision. We conduct extensive experimental analysis and ablation studies. Comprehensive evaluations on nearly 20 benchmarks show that Thyme yields significant and consistent performance gains, particularly in challenging high-resolution perception and complex reasoning tasks.
title Thyme: Think Beyond Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.11630