LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shuai, Zhang, Daoan, Bai, Tianyi, Shao, Shitong, Luo, Jiebo, Wei, Jiaheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918216550318080
author Wang, Shuai
Zhang, Daoan
Bai, Tianyi
Shao, Shitong
Luo, Jiebo
Wei, Jiaheng
author_facet Wang, Shuai
Zhang, Daoan
Bai, Tianyi
Shao, Shitong
Luo, Jiebo
Wei, Jiaheng
contents Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-of-the-art VLMs still struggle to understand 3D space and long videos, although they are powerful in typical vision-language tasks. Current methods often rely on specialized architectural designs to improve performance for 3D tasks and video understanding tasks separately. In contrast, we propose LAST, short for LeArn to Think in Space and Time, to jointly improve 3D spatial and long video understanding for general VLMs with only a set of 2D images as inputs. LAST makes VLMs think in space and time rather than only with text before giving the final answer, building visual thinking trajectories in 3D space and temporal dimension. We demonstrate the effectiveness of LAST in two scenarios: 1) zero-shot, where we directly prompt proprietary models; and 2) fine-tuning general VLMs with data that include thinking trajectories in 3D space and time. We show that LAST brings substantial gains in various benchmarks, including 3 spatial understanding, 4 video understanding, and 3 image understanding tasks. Notably, 15.8% gains on EgoSchema with GPT-4o in a zero-shot manner and 8.3 gains on VSI-Bench compared with Qwen2.5-VL-7B.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19261
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
Wang, Shuai
Zhang, Daoan
Bai, Tianyi
Shao, Shitong
Luo, Jiebo
Wei, Jiaheng
Computer Vision and Pattern Recognition
Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-of-the-art VLMs still struggle to understand 3D space and long videos, although they are powerful in typical vision-language tasks. Current methods often rely on specialized architectural designs to improve performance for 3D tasks and video understanding tasks separately. In contrast, we propose LAST, short for LeArn to Think in Space and Time, to jointly improve 3D spatial and long video understanding for general VLMs with only a set of 2D images as inputs. LAST makes VLMs think in space and time rather than only with text before giving the final answer, building visual thinking trajectories in 3D space and temporal dimension. We demonstrate the effectiveness of LAST in two scenarios: 1) zero-shot, where we directly prompt proprietary models; and 2) fine-tuning general VLMs with data that include thinking trajectories in 3D space and time. We show that LAST brings substantial gains in various benchmarks, including 3 spatial understanding, 4 video understanding, and 3 image understanding tasks. Notably, 15.8% gains on EgoSchema with GPT-4o in a zero-shot manner and 8.3 gains on VSI-Bench compared with Qwen2.5-VL-7B.
title LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.19261