LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jiangyong, Ma, Xiaojian, Linghu, Xiongkun, He, Junchao, Li, Qing, Zhu, Song-Chun, Chen, Yixin, Jia, Baoxiong, Huang, Siyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917357228654592
author Huang, Jiangyong
Ma, Xiaojian
Linghu, Xiongkun
He, Junchao
Li, Qing
Zhu, Song-Chun
Chen, Yixin
Jia, Baoxiong
Huang, Siyuan
author_facet Huang, Jiangyong
Ma, Xiaojian
Linghu, Xiongkun
He, Junchao
Li, Qing
Zhu, Song-Chun
Chen, Yixin
Jia, Baoxiong
Huang, Siyuan
contents Developing vision-language models (VLMs) capable of understanding 3D scenes has been a longstanding research goal. Despite recent progress, 3D VLMs still struggle with spatial reasoning and robustness. We identify three key obstacles hindering their progress: (1) scene representation is constrained by a capacity-efficiency trade-off, which impedes scalable learning; (2) training data lacks a comprehensive scheme, with limited diversity across tasks and scene domains; and (3) models exhibit robustness deficiencies and lack effective post-training. To address these challenges, we first propose condensed feature grid (CFG), an efficient scene representation that significantly reduces token overhead while preserving strong perceptual capacity. Building on CFG, we introduce LEO-VL, a 3D VLM trained on over 700k 3D vision-language (3D-VL) data spanning four real-world indoor domains and five tasks such as captioning and dialogue. To further improve robustness, we propose SceneDPO, a novel post-training objective that incorporates contrastive signals across both answers and scenes. LEO-VL achieves state-of-the-art performance on various 3D-VL benchmarks, such as SQA3D, Beacon3D, and Scan2Cap. Extensive analyses highlight the efficiency of CFG and provide key insights such as the importance of task and scene diversity, the priority of data quality for effective scaling, and the advantages of SceneDPO.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09935
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
Huang, Jiangyong
Ma, Xiaojian
Linghu, Xiongkun
He, Junchao
Li, Qing
Zhu, Song-Chun
Chen, Yixin
Jia, Baoxiong
Huang, Siyuan
Computer Vision and Pattern Recognition
Developing vision-language models (VLMs) capable of understanding 3D scenes has been a longstanding research goal. Despite recent progress, 3D VLMs still struggle with spatial reasoning and robustness. We identify three key obstacles hindering their progress: (1) scene representation is constrained by a capacity-efficiency trade-off, which impedes scalable learning; (2) training data lacks a comprehensive scheme, with limited diversity across tasks and scene domains; and (3) models exhibit robustness deficiencies and lack effective post-training. To address these challenges, we first propose condensed feature grid (CFG), an efficient scene representation that significantly reduces token overhead while preserving strong perceptual capacity. Building on CFG, we introduce LEO-VL, a 3D VLM trained on over 700k 3D vision-language (3D-VL) data spanning four real-world indoor domains and five tasks such as captioning and dialogue. To further improve robustness, we propose SceneDPO, a novel post-training objective that incorporates contrastive signals across both answers and scenes. LEO-VL achieves state-of-the-art performance on various 3D-VL benchmarks, such as SQA3D, Beacon3D, and Scan2Cap. Extensive analyses highlight the efficiency of CFG and provide key insights such as the importance of task and scene diversity, the priority of data quality for effective scaling, and the advantages of SceneDPO.
title LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.09935