Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qin, Luozheng, Gong, Jia, Sun, Yuqing, Li, Tianjiao, Yang, Mengping, Yang, Xiaomeng, Qu, Chao, Tan, Zhiyu, Li, Hao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917301044903936
author Qin, Luozheng
Gong, Jia
Sun, Yuqing
Li, Tianjiao
Yang, Mengping
Yang, Xiaomeng
Qu, Chao
Tan, Zhiyu
Li, Hao
author_facet Qin, Luozheng
Gong, Jia
Sun, Yuqing
Li, Tianjiao
Yang, Mengping
Yang, Xiaomeng
Qu, Chao
Tan, Zhiyu
Li, Hao
contents Chain-of-Thought (CoT) reasoning has been widely adopted to enhance Large Language Models (LLMs) by decomposing complex tasks into simpler, sequential subtasks. However, extending CoT to vision-language reasoning tasks remains challenging, as it often requires interpreting transitions of visual states to support reasoning. Existing methods often struggle with this due to limited capacity of modeling visual state transitions or incoherent visual trajectories caused by fragmented architectures. To overcome these limitations, we propose Uni-CoT, a Unified Chain-of-Thought framework that enables coherent and grounded multimodal reasoning within a single unified model. The key idea is to leverage a model capable of both image understanding and generation to reason over visual content and model evolving visual states. However, empowering a unified model to achieve that is non-trivial, given the high computational cost and the burden of training. To address this, Uni-CoT introduces a novel two-level reasoning paradigm: A Macro-Level CoT for high-level task planning and A Micro-Level CoT for subtask execution. This design significantly reduces the computational overhead. Furthermore, we introduce a structured training paradigm that combines interleaved image-text supervision for macro-level CoT with multi-task objectives for micro-level CoT. Together, these innovations allow Uni-CoT to perform scalable and coherent multi-modal reasoning. Furthermore, thanks to our design, all experiments can be efficiently completed using only 8 A100 GPUs with 80GB VRAM each. Experimental results on reasoning-driven image generation benchmark (WISE) and editing benchmarks (RISE and KRIS) indicates that Uni-CoT demonstrates SOTA performance and strong generalization, establishing Uni-CoT as a promising solution for multi-modal reasoning. Project Page and Code: https://sais-fuxi.github.io/projects/uni-cot/
format Preprint
id arxiv_https___arxiv_org_abs_2508_05606
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
Qin, Luozheng
Gong, Jia
Sun, Yuqing
Li, Tianjiao
Yang, Mengping
Yang, Xiaomeng
Qu, Chao
Tan, Zhiyu
Li, Hao
Computer Vision and Pattern Recognition
Computation and Language
Chain-of-Thought (CoT) reasoning has been widely adopted to enhance Large Language Models (LLMs) by decomposing complex tasks into simpler, sequential subtasks. However, extending CoT to vision-language reasoning tasks remains challenging, as it often requires interpreting transitions of visual states to support reasoning. Existing methods often struggle with this due to limited capacity of modeling visual state transitions or incoherent visual trajectories caused by fragmented architectures. To overcome these limitations, we propose Uni-CoT, a Unified Chain-of-Thought framework that enables coherent and grounded multimodal reasoning within a single unified model. The key idea is to leverage a model capable of both image understanding and generation to reason over visual content and model evolving visual states. However, empowering a unified model to achieve that is non-trivial, given the high computational cost and the burden of training. To address this, Uni-CoT introduces a novel two-level reasoning paradigm: A Macro-Level CoT for high-level task planning and A Micro-Level CoT for subtask execution. This design significantly reduces the computational overhead. Furthermore, we introduce a structured training paradigm that combines interleaved image-text supervision for macro-level CoT with multi-task objectives for micro-level CoT. Together, these innovations allow Uni-CoT to perform scalable and coherent multi-modal reasoning. Furthermore, thanks to our design, all experiments can be efficiently completed using only 8 A100 GPUs with 80GB VRAM each. Experimental results on reasoning-driven image generation benchmark (WISE) and editing benchmarks (RISE and KRIS) indicates that Uni-CoT demonstrates SOTA performance and strong generalization, establishing Uni-CoT as a promising solution for multi-modal reasoning. Project Page and Code: https://sais-fuxi.github.io/projects/uni-cot/
title Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2508.05606