Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Guanhua, Yao, Yutong, Sun, Shenghe, Gao, Ci-Jun, Liu, Shudong, Chao, Lidia S., Wan, Feng, Wong, Derek F.
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914567527858176
author Chen, Guanhua
Yao, Yutong
Sun, Shenghe
Gao, Ci-Jun
Liu, Shudong
Chao, Lidia S.
Wan, Feng
Wong, Derek F.
author_facet Chen, Guanhua
Yao, Yutong
Sun, Shenghe
Gao, Ci-Jun
Liu, Shudong
Chao, Lidia S.
Wan, Feng
Wong, Derek F.
contents Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure question answering (VP-QA) remains largely unexplored. VP-QA presents unique challenges where users query next-step actions by uploading images for intermediate states of complex procedures. To systematically evaluate VLMs on this practical task, we propose ProcedureVQA, a novel multimodal benchmark specifically designed for visual procedural reasoning. Through comprehensive analysis, we identify two critical limitations in current VLMs: inadequate cross-modal retrieval of structured procedures given visual states, and misalignment between image sequence granularity and textual step decomposition. To address these issues, we present Chain-of-Procedure (CoP), a hierarchical reasoning framework that first retrieves relevant instructions using visual cues, then performs step refinement through semantic decomposition, and finally generates the next step. Experiments across six VLMs demonstrate CoP's effectiveness, achieving up to 13% absolute improvement over standard baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14928
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
Chen, Guanhua
Yao, Yutong
Sun, Shenghe
Gao, Ci-Jun
Liu, Shudong
Chao, Lidia S.
Wan, Feng
Wong, Derek F.
Computation and Language
Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure question answering (VP-QA) remains largely unexplored. VP-QA presents unique challenges where users query next-step actions by uploading images for intermediate states of complex procedures. To systematically evaluate VLMs on this practical task, we propose ProcedureVQA, a novel multimodal benchmark specifically designed for visual procedural reasoning. Through comprehensive analysis, we identify two critical limitations in current VLMs: inadequate cross-modal retrieval of structured procedures given visual states, and misalignment between image sequence granularity and textual step decomposition. To address these issues, we present Chain-of-Procedure (CoP), a hierarchical reasoning framework that first retrieves relevant instructions using visual cues, then performs step refinement through semantic decomposition, and finally generates the next step. Experiments across six VLMs demonstrate CoP's effectiveness, achieving up to 13% absolute improvement over standard baselines.
title Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
topic Computation and Language
url https://arxiv.org/abs/2605.14928