ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Hasegawa, Kimihiro, Imrattanatrai, Wiradee, Cheng, Zhi-Qi, Asada, Masaki, Holm, Susan, Wang, Yuran, Fukuda, Ken, Mitamura, Teruko |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
by: Hasegawa, Kimihiro, et al.
Published: (2025)
by: Hasegawa, Kimihiro, et al.
Published: (2025)
TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
by: Hasegawa, Kimihiro, et al.
Published: (2025)
by: Hasegawa, Kimihiro, et al.
Published: (2025)
A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
by: Imrattanatrai, Wiradee, et al.
Published: (2025)
by: Imrattanatrai, Wiradee, et al.
Published: (2025)
Formulation Comparison for Timeline Construction using LLMs
by: Hasegawa, Kimihiro, et al.
Published: (2024)
by: Hasegawa, Kimihiro, et al.
Published: (2024)
HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
by: Egami, Shusaku, et al.
Published: (2026)
by: Egami, Shusaku, et al.
Published: (2026)
ColBERT Retrieval and Ensemble Response Scoring for Language Model Question Answering
by: Gichamba, Alex, et al.
Published: (2024)
by: Gichamba, Alex, et al.
Published: (2024)
Scaling Multi-Document Event Summarization: Evaluating Compression vs. Full-Text Approaches
by: Pratapa, Adithya, et al.
Published: (2025)
by: Pratapa, Adithya, et al.
Published: (2025)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
by: Ikoma, Hayato, et al.
Published: (2025)
by: Ikoma, Hayato, et al.
Published: (2025)
Estimating Optimal Context Length for Hybrid Retrieval-augmented Multi-document Summarization
by: Pratapa, Adithya, et al.
Published: (2025)
by: Pratapa, Adithya, et al.
Published: (2025)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark
by: Liang, Zhenwen, et al.
Published: (2024)
by: Liang, Zhenwen, et al.
Published: (2024)
PokeMQA: Programmable knowledge editing for Multi-hop Question Answering
by: Gu, Hengrui, et al.
Published: (2023)
by: Gu, Hengrui, et al.
Published: (2023)
PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
by: Bahaj, Adil, et al.
Published: (2025)
by: Bahaj, Adil, et al.
Published: (2025)
Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
by: Kong, Yaxuan, et al.
Published: (2025)
by: Kong, Yaxuan, et al.
Published: (2025)
BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions
by: Sengupta, Saptarshi, et al.
Published: (2025)
by: Sengupta, Saptarshi, et al.
Published: (2025)
MQA-KEAL: Multi-hop Question Answering under Knowledge Editing for Arabic Language
by: Ali, Muhammad Asif, et al.
Published: (2024)
by: Ali, Muhammad Asif, et al.
Published: (2024)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
by: Bertsch, Amanda, et al.
Published: (2025)
by: Bertsch, Amanda, et al.
Published: (2025)
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
by: Lim, Qi Zhi, et al.
Published: (2025)
by: Lim, Qi Zhi, et al.
Published: (2025)
Graph Guided Question Answer Generation for Procedural Question-Answering
by: Pham, Hai X., et al.
Published: (2024)
by: Pham, Hai X., et al.
Published: (2024)
SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
by: Pramanick, Shraman, et al.
Published: (2024)
by: Pramanick, Shraman, et al.
Published: (2024)
Understanding and Supporting Formal Email Exchange by Answering AI-Generated Questions
by: Miura, Yusuke, et al.
Published: (2025)
by: Miura, Yusuke, et al.
Published: (2025)
Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
by: Li, Heng, et al.
Published: (2024)
by: Li, Heng, et al.
Published: (2024)
The State of money in circulation and its reform in Konbaung Burma : The 1790s–1860s / Teruko Saito
by: Saito, Teruko
by: Saito, Teruko
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
by: Yu, Zeping, et al.
Published: (2024)
by: Yu, Zeping, et al.
Published: (2024)
ProCQA: A Large-scale Community-based Programming Question Answering Dataset for Code Search
by: Li, Zehan, et al.
Published: (2024)
by: Li, Zehan, et al.
Published: (2024)
English-to-Japanese Cross-Language Question-Answering System using Weighted Adding with Multiple Answers
by: Masaki Murata
Published: (2009)
by: Masaki Murata
Published: (2009)
Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments
by: Ugai, Takanori, et al.
Published: (2024)
by: Ugai, Takanori, et al.
Published: (2024)
PAQA: Toward ProActive Open-Retrieval Question Answering
by: Erbacher, Pierre, et al.
Published: (2024)
by: Erbacher, Pierre, et al.
Published: (2024)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
by: Fernando, Basura, et al.
Published: (2025)
by: Fernando, Basura, et al.
Published: (2025)
Causal Understanding For Video Question Answering
by: Guda, Bhanu Prakash Reddy, et al.
Published: (2024)
by: Guda, Bhanu Prakash Reddy, et al.
Published: (2024)
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
by: Peddi, Rohith, et al.
Published: (2023)
by: Peddi, Rohith, et al.
Published: (2023)
Multilingual Open QA on the MIA Shared Task
by: Yarrabelly, Navya, et al.
Published: (2025)
by: Yarrabelly, Navya, et al.
Published: (2025)
RoadscapesQA: A Multitask, Multimodal Dataset for Visual Question Answering on Indian Roads
by: Iyer, Vijayasri, et al.
Published: (2026)
by: Iyer, Vijayasri, et al.
Published: (2026)
Modelos teóricos atuais da dislexia do desenvolvimento
by: Olinda Teruko Kajihara
Published: (2008)
by: Olinda Teruko Kajihara
Published: (2008)
PolQA: Polish Question Answering Dataset
by: Rybak, Piotr, et al.
Published: (2022)
by: Rybak, Piotr, et al.
Published: (2022)
A Collection of Question Answering Datasets for Norwegian
by: Mikhailov, Vladislav, et al.
Published: (2025)
by: Mikhailov, Vladislav, et al.
Published: (2025)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
by: Sood, Ekta, et al.
Published: (2021)
by: Sood, Ekta, et al.
Published: (2021)
Multimodal Rationales for Explainable Visual Question Answering
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Memory-Centric Embodied Question Answering
by: Zhai, Mingliang, et al.
Published: (2025)
by: Zhai, Mingliang, et al.
Published: (2025)
Similar Items
-
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
by: Hasegawa, Kimihiro, et al.
Published: (2025) -
TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
by: Hasegawa, Kimihiro, et al.
Published: (2025) -
A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
by: Imrattanatrai, Wiradee, et al.
Published: (2025) -
Formulation Comparison for Timeline Construction using LLMs
by: Hasegawa, Kimihiro, et al.
Published: (2024) -
HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
by: Egami, Shusaku, et al.
Published: (2026)