HCQA-1.5 @ Ego4D EgoSchema Challenge 2025

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Haoyu, Feng, Yisen, Chu, Qiaohui, Liu, Meng, Guan, Weili, Wang, Yaowei, Nie, Liqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910969776570368
author Zhang, Haoyu
Feng, Yisen
Chu, Qiaohui
Liu, Meng
Guan, Weili
Wang, Yaowei
Nie, Liqiang
author_facet Zhang, Haoyu
Feng, Yisen
Chu, Qiaohui
Liu, Meng
Guan, Weili
Wang, Yaowei
Nie, Liqiang
contents In this report, we present the method that achieves third place for Ego4D EgoSchema Challenge in CVPR 2025. To improve the reliability of answer prediction in egocentric video question answering, we propose an effective extension to the previously proposed HCQA framework. Our approach introduces a multi-source aggregation strategy to generate diverse predictions, followed by a confidence-based filtering mechanism that selects high-confidence answers directly. For low-confidence cases, we incorporate a fine-grained reasoning module that performs additional visual and contextual analysis to refine the predictions. Evaluated on the EgoSchema blind test set, our method achieves 77% accuracy on over 5,000 human-curated multiple-choice questions, outperforming last year's winning solution and the majority of participating teams. Our code will be added at https://github.com/Hyu-Zhang/HCQA.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20644
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
Zhang, Haoyu
Feng, Yisen
Chu, Qiaohui
Liu, Meng
Guan, Weili
Wang, Yaowei
Nie, Liqiang
Computer Vision and Pattern Recognition
Artificial Intelligence
In this report, we present the method that achieves third place for Ego4D EgoSchema Challenge in CVPR 2025. To improve the reliability of answer prediction in egocentric video question answering, we propose an effective extension to the previously proposed HCQA framework. Our approach introduces a multi-source aggregation strategy to generate diverse predictions, followed by a confidence-based filtering mechanism that selects high-confidence answers directly. For low-confidence cases, we incorporate a fine-grained reasoning module that performs additional visual and contextual analysis to refine the predictions. Evaluated on the EgoSchema blind test set, our method achieves 77% accuracy on over 5,000 human-curated multiple-choice questions, outperforming last year's winning solution and the majority of participating teams. Our code will be added at https://github.com/Hyu-Zhang/HCQA.
title HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.20644