Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Guanhua, Yao, Yutong, Sun, Shenghe, Gao, Ci-Jun, Liu, Shudong, Chao, Lidia S., Wan, Feng, Wong, Derek F.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914567527858176
author Chen, Guanhua
Yao, Yutong
Sun, Shenghe
Gao, Ci-Jun
Liu, Shudong
Chao, Lidia S.
Wan, Feng
Wong, Derek F.
author_facet Chen, Guanhua
Yao, Yutong
Sun, Shenghe
Gao, Ci-Jun
Liu, Shudong
Chao, Lidia S.
Wan, Feng
Wong, Derek F.
contents Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure question answering (VP-QA) remains largely unexplored. VP-QA presents unique challenges where users query next-step actions by uploading images for intermediate states of complex procedures. To systematically evaluate VLMs on this practical task, we propose ProcedureVQA, a novel multimodal benchmark specifically designed for visual procedural reasoning. Through comprehensive analysis, we identify two critical limitations in current VLMs: inadequate cross-modal retrieval of structured procedures given visual states, and misalignment between image sequence granularity and textual step decomposition. To address these issues, we present Chain-of-Procedure (CoP), a hierarchical reasoning framework that first retrieves relevant instructions using visual cues, then performs step refinement through semantic decomposition, and finally generates the next step. Experiments across six VLMs demonstrate CoP's effectiveness, achieving up to 13% absolute improvement over standard baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14928
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
Chen, Guanhua
Yao, Yutong
Sun, Shenghe
Gao, Ci-Jun
Liu, Shudong
Chao, Lidia S.
Wan, Feng
Wong, Derek F.
Computation and Language
Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure question answering (VP-QA) remains largely unexplored. VP-QA presents unique challenges where users query next-step actions by uploading images for intermediate states of complex procedures. To systematically evaluate VLMs on this practical task, we propose ProcedureVQA, a novel multimodal benchmark specifically designed for visual procedural reasoning. Through comprehensive analysis, we identify two critical limitations in current VLMs: inadequate cross-modal retrieval of structured procedures given visual states, and misalignment between image sequence granularity and textual step decomposition. To address these issues, we present Chain-of-Procedure (CoP), a hierarchical reasoning framework that first retrieves relevant instructions using visual cues, then performs step refinement through semantic decomposition, and finally generates the next step. Experiments across six VLMs demonstrate CoP's effectiveness, achieving up to 13% absolute improvement over standard baselines.
title Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
topic Computation and Language
url https://arxiv.org/abs/2605.14928