STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chen, Zhang, Han, Yang, Zhantao, Chen, Fangyi, Wang, Zihan, Bolimera, Anudeepsekhar, Savvides, Marios
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915787682349056
author Li, Chen
Zhang, Han
Yang, Zhantao
Chen, Fangyi
Wang, Zihan
Bolimera, Anudeepsekhar
Savvides, Marios
author_facet Li, Chen
Zhang, Han
Yang, Zhantao
Chen, Fangyi
Wang, Zihan
Bolimera, Anudeepsekhar
Savvides, Marios
contents Vision-language models (VLMs) have made significant strides in reasoning, yet they often struggle with complex multimodal tasks and tend to generate overly verbose outputs. A key limitation is their reliance on chain-of-thought (CoT) reasoning, despite many tasks benefiting from alternative topologies like trees or graphs. To address this, we introduce STELAR-Vision, a training framework for topology-aware reasoning. At its core is TopoAug, a synthetic data pipeline that enriches training with diverse topological structures. Using supervised fine-tuning and reinforcement learning, we post-train Qwen2VL models with both accuracy and efficiency in mind. Additionally, we propose Frugal Learning, which reduces output length with minimal accuracy loss. On MATH-V and VLM-S2H, STELAR-Vision improves accuracy by 9.7% over its base model and surpasses the larger Qwen2VL-72B-Instruct by 7.3%. On five out-of-distribution benchmarks, it outperforms Phi-4-Multimodal-Instruct by up to 28.4% and LLaMA-3.2-11B-Vision-Instruct by up to 13.2%, demonstrating strong generalization. Compared to Chain-Only training, our approach achieves 4.3% higher overall accuracy on in-distribution datasets and consistently outperforms across all OOD benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08688
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
Li, Chen
Zhang, Han
Yang, Zhantao
Chen, Fangyi
Wang, Zihan
Bolimera, Anudeepsekhar
Savvides, Marios
Artificial Intelligence
Computer Vision and Pattern Recognition
Vision-language models (VLMs) have made significant strides in reasoning, yet they often struggle with complex multimodal tasks and tend to generate overly verbose outputs. A key limitation is their reliance on chain-of-thought (CoT) reasoning, despite many tasks benefiting from alternative topologies like trees or graphs. To address this, we introduce STELAR-Vision, a training framework for topology-aware reasoning. At its core is TopoAug, a synthetic data pipeline that enriches training with diverse topological structures. Using supervised fine-tuning and reinforcement learning, we post-train Qwen2VL models with both accuracy and efficiency in mind. Additionally, we propose Frugal Learning, which reduces output length with minimal accuracy loss. On MATH-V and VLM-S2H, STELAR-Vision improves accuracy by 9.7% over its base model and surpasses the larger Qwen2VL-72B-Instruct by 7.3%. On five out-of-distribution benchmarks, it outperforms Phi-4-Multimodal-Instruct by up to 28.4% and LLaMA-3.2-11B-Vision-Instruct by up to 13.2%, demonstrating strong generalization. Compared to Chain-Only training, our approach achieves 4.3% higher overall accuracy on in-distribution datasets and consistently outperforms across all OOD benchmarks.
title STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.08688