IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Chen, Sugandhika, Chinthani, Ee, Yeo Keat, Peh, Eric, Zhang, Hao, Yang, Hong, Rajan, Deepu, Fernando, Basura
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918113508851712
author Li, Chen
Sugandhika, Chinthani
Ee, Yeo Keat
Peh, Eric
Zhang, Hao
Yang, Hong
Rajan, Deepu
Fernando, Basura
author_facet Li, Chen
Sugandhika, Chinthani
Ee, Yeo Keat
Peh, Eric
Zhang, Hao
Yang, Hong
Rajan, Deepu
Fernando, Basura
contents Existing human motion Q\&A methods rely on explicit program execution, where the requirement for manually defined functional modules may limit the scalability and adaptability. To overcome this, we propose an implicit program-guided motion reasoning (IMoRe) framework that unifies reasoning across multiple query types without manually designed modules. Unlike existing implicit reasoning approaches that infer reasoning operations from question words, our model directly conditions on structured program functions, ensuring a more precise execution of reasoning steps. Additionally, we introduce a program-guided reading mechanism, which dynamically selects multi-level motion representations from a pretrained motion Vision Transformer (ViT), capturing both high-level semantics and fine-grained motion cues. The reasoning module iteratively refines memory representations, leveraging structured program functions to extract relevant information for different query types. Our model achieves state-of-the-art performance on Babel-QA and generalizes to a newly constructed motion Q\&A dataset based on HuMMan, demonstrating its adaptability across different motion reasoning datasets. Code and dataset are available at: https://github.com/LUNAProject22/IMoRe.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01984
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A
Li, Chen
Sugandhika, Chinthani
Ee, Yeo Keat
Peh, Eric
Zhang, Hao
Yang, Hong
Rajan, Deepu
Fernando, Basura
Computer Vision and Pattern Recognition
Existing human motion Q\&A methods rely on explicit program execution, where the requirement for manually defined functional modules may limit the scalability and adaptability. To overcome this, we propose an implicit program-guided motion reasoning (IMoRe) framework that unifies reasoning across multiple query types without manually designed modules. Unlike existing implicit reasoning approaches that infer reasoning operations from question words, our model directly conditions on structured program functions, ensuring a more precise execution of reasoning steps. Additionally, we introduce a program-guided reading mechanism, which dynamically selects multi-level motion representations from a pretrained motion Vision Transformer (ViT), capturing both high-level semantics and fine-grained motion cues. The reasoning module iteratively refines memory representations, leveraging structured program functions to extract relevant information for different query types. Our model achieves state-of-the-art performance on Babel-QA and generalizes to a newly constructed motion Q\&A dataset based on HuMMan, demonstrating its adaptability across different motion reasoning datasets. Code and dataset are available at: https://github.com/LUNAProject22/IMoRe.
title IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.01984