SpatialTraceGen: High-Fidelity Traces for Efficient VLM Spatial Reasoning Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huh, Gio, Sheth, Dhruv, Zirvi, Rayhan, Xiao, Frank
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909880783208448
author Huh, Gio
Sheth, Dhruv
Zirvi, Rayhan
Xiao, Frank
author_facet Huh, Gio
Sheth, Dhruv
Zirvi, Rayhan
Xiao, Frank
contents While Vision-Language Models (VLMs) excel in many areas, they struggle with complex spatial reasoning, which requires problem decomposition and strategic tool use. Fine-tuning smaller, more deployable models offers an efficient path to strong performance, but this is hampered by a major bottleneck: the absence of high-quality, step-by-step reasoning data. To address this data-efficiency gap, we introduce SpatialTraceGen, a framework to distill the reasoning processes of a large teacher model into a high-quality dataset of multi-hop, multi-tool reasoning traces. A key innovation is our automated Verifier, which scalably ensures the fidelity of each reasoning step, providing a cost-effective alternative to manual human annotation. On the CLEVR-Humans benchmark, this verifier-guided process improves the average quality score of traces by 17\% while reducing quality variance by over 40\%. SpatialTraceGen delivers a dataset of expert traces, providing the structured, step-by-step examples of tool use necessary for effective fine-tuning and sample-efficient offline reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00054
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpatialTraceGen: High-Fidelity Traces for Efficient VLM Spatial Reasoning Distillation
Huh, Gio
Sheth, Dhruv
Zirvi, Rayhan
Xiao, Frank
Machine Learning
Artificial Intelligence
While Vision-Language Models (VLMs) excel in many areas, they struggle with complex spatial reasoning, which requires problem decomposition and strategic tool use. Fine-tuning smaller, more deployable models offers an efficient path to strong performance, but this is hampered by a major bottleneck: the absence of high-quality, step-by-step reasoning data. To address this data-efficiency gap, we introduce SpatialTraceGen, a framework to distill the reasoning processes of a large teacher model into a high-quality dataset of multi-hop, multi-tool reasoning traces. A key innovation is our automated Verifier, which scalably ensures the fidelity of each reasoning step, providing a cost-effective alternative to manual human annotation. On the CLEVR-Humans benchmark, this verifier-guided process improves the average quality score of traces by 17\% while reducing quality variance by over 40\%. SpatialTraceGen delivers a dataset of expert traces, providing the structured, step-by-step examples of tool use necessary for effective fine-tuning and sample-efficient offline reinforcement learning.
title SpatialTraceGen: High-Fidelity Traces for Efficient VLM Spatial Reasoning Distillation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.00054