Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Sicheng, Huang, Yukai, Sun, Shitong, Cai, Weitong, Deng, Jiankang, Song, Jifei, Zhang, Zhensong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911377267884032
author Yang, Sicheng
Huang, Yukai
Sun, Shitong
Cai, Weitong
Deng, Jiankang
Song, Jifei
Zhang, Zhensong
author_facet Yang, Sicheng
Huang, Yukai
Sun, Shitong
Cai, Weitong
Deng, Jiankang
Song, Jifei
Zhang, Zhensong
contents Multimodal Large Language Models (MLLMs) struggle with complex video QA benchmarks like HD-EPIC VQA due to ambiguous queries/options, poor long-range temporal reasoning, and non-standardized outputs. We propose a framework integrating query/choice pre-processing, domain-specific Qwen2.5-VL fine-tuning, a novel Temporal Chain-of-Thought (T-CoT) prompting for multi-step reasoning, and robust post-processing. This system achieves 41.6% accuracy on HD-EPIC VQA, highlighting the need for holistic pipeline optimization in demanding video understanding. Our code, fine-tuned models are available at https://github.com/YoungSeng/Egocentric-Co-Pilot.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10228
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
Yang, Sicheng
Huang, Yukai
Sun, Shitong
Cai, Weitong
Deng, Jiankang
Song, Jifei
Zhang, Zhensong
Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
Multimodal Large Language Models (MLLMs) struggle with complex video QA benchmarks like HD-EPIC VQA due to ambiguous queries/options, poor long-range temporal reasoning, and non-standardized outputs. We propose a framework integrating query/choice pre-processing, domain-specific Qwen2.5-VL fine-tuning, a novel Temporal Chain-of-Thought (T-CoT) prompting for multi-step reasoning, and robust post-processing. This system achieves 41.6% accuracy on HD-EPIC VQA, highlighting the need for holistic pipeline optimization in demanding video understanding. Our code, fine-tuned models are available at https://github.com/YoungSeng/Egocentric-Co-Pilot.
title Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
topic Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
url https://arxiv.org/abs/2601.10228