Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914567527858176 |
|---|---|
| author | Chen, Guanhua Yao, Yutong Sun, Shenghe Gao, Ci-Jun Liu, Shudong Chao, Lidia S. Wan, Feng Wong, Derek F. |
| author_facet | Chen, Guanhua Yao, Yutong Sun, Shenghe Gao, Ci-Jun Liu, Shudong Chao, Lidia S. Wan, Feng Wong, Derek F. |
| contents | Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure question answering (VP-QA) remains largely unexplored. VP-QA presents unique challenges where users query next-step actions by uploading images for intermediate states of complex procedures. To systematically evaluate VLMs on this practical task, we propose ProcedureVQA, a novel multimodal benchmark specifically designed for visual procedural reasoning. Through comprehensive analysis, we identify two critical limitations in current VLMs: inadequate cross-modal retrieval of structured procedures given visual states, and misalignment between image sequence granularity and textual step decomposition. To address these issues, we present Chain-of-Procedure (CoP), a hierarchical reasoning framework that first retrieves relevant instructions using visual cues, then performs step refinement through semantic decomposition, and finally generates the next step. Experiments across six VLMs demonstrate CoP's effectiveness, achieving up to 13% absolute improvement over standard baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_14928 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA Chen, Guanhua Yao, Yutong Sun, Shenghe Gao, Ci-Jun Liu, Shudong Chao, Lidia S. Wan, Feng Wong, Derek F. Computation and Language Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure question answering (VP-QA) remains largely unexplored. VP-QA presents unique challenges where users query next-step actions by uploading images for intermediate states of complex procedures. To systematically evaluate VLMs on this practical task, we propose ProcedureVQA, a novel multimodal benchmark specifically designed for visual procedural reasoning. Through comprehensive analysis, we identify two critical limitations in current VLMs: inadequate cross-modal retrieval of structured procedures given visual states, and misalignment between image sequence granularity and textual step decomposition. To address these issues, we present Chain-of-Procedure (CoP), a hierarchical reasoning framework that first retrieves relevant instructions using visual cues, then performs step refinement through semantic decomposition, and finally generates the next step. Experiments across six VLMs demonstrate CoP's effectiveness, achieving up to 13% absolute improvement over standard baselines. |
| title | Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2605.14928 |