A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kurpath, Mohammed Irfan, Kaithakkodan, Jaseel Muhammad, Zhou, Jinxing, Mullappilly, Sahal Shaji, Almansoori, Mohammad, Ahsan, Noor, Kalmakhanbet, Beknur, Shikhar, Sambal, Lalla, Rishabh, Lahoud, Jean, Awad, Mariette, Khan, Fahad Shahbaz, Khan, Salman, Anwer, Rao Muhammad, Cholakkal, Hisham
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912775796686848
author Kurpath, Mohammed Irfan
Kaithakkodan, Jaseel Muhammad
Zhou, Jinxing
Mullappilly, Sahal Shaji
Almansoori, Mohammad
Ahsan, Noor
Kalmakhanbet, Beknur
Shikhar, Sambal
Lalla, Rishabh
Lahoud, Jean
Awad, Mariette
Khan, Fahad Shahbaz
Khan, Salman
Anwer, Rao Muhammad
Cholakkal, Hisham
author_facet Kurpath, Mohammed Irfan
Kaithakkodan, Jaseel Muhammad
Zhou, Jinxing
Mullappilly, Sahal Shaji
Almansoori, Mohammad
Ahsan, Noor
Kalmakhanbet, Beknur
Shikhar, Sambal
Lalla, Rishabh
Lahoud, Jean
Awad, Mariette
Khan, Fahad Shahbaz
Khan, Salman
Anwer, Rao Muhammad
Cholakkal, Hisham
contents Long-form multimodal video understanding requires integrating vision, speech, and ambient audio with coherent long-range reasoning. Existing benchmarks emphasize either temporal length or multimodal richness, but rarely both and while some incorporate open-ended questions and advanced metrics, they mostly rely on single-score accuracy, obscuring failure modes. We introduce LongShOTBench, a diagnostic benchmark with open-ended, intent-driven questions; single- and multi-turn dialogues; and tasks requiring multimodal reasoning and agentic tool use across video, audio, and speech. Each item includes a reference answer and graded rubric for interpretable, and traceable evaluation. LongShOTBench is produced via a scalable, human-validated pipeline to ensure coverage and reproducibility. All samples in our LongShOTBench are human-verified and corrected. Furthermore, we present LongShOTAgent, an agentic system that analyzes long videos via preprocessing, search, and iterative refinement. On LongShOTBench, state-of-the-art MLLMs show large gaps: Gemini-2.5-Flash achieves 52.95%, open-source models remain below 30%, and LongShOTAgent attains 44.66%. These results underscore the difficulty of real-world long-form video understanding. LongShOTBench provides a practical, reproducible foundation for evaluating and improving MLLMs. All resources are available on GitHub: https://github.com/mbzuai-oryx/longshot.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16978
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
Kurpath, Mohammed Irfan
Kaithakkodan, Jaseel Muhammad
Zhou, Jinxing
Mullappilly, Sahal Shaji
Almansoori, Mohammad
Ahsan, Noor
Kalmakhanbet, Beknur
Shikhar, Sambal
Lalla, Rishabh
Lahoud, Jean
Awad, Mariette
Khan, Fahad Shahbaz
Khan, Salman
Anwer, Rao Muhammad
Cholakkal, Hisham
Computer Vision and Pattern Recognition
Long-form multimodal video understanding requires integrating vision, speech, and ambient audio with coherent long-range reasoning. Existing benchmarks emphasize either temporal length or multimodal richness, but rarely both and while some incorporate open-ended questions and advanced metrics, they mostly rely on single-score accuracy, obscuring failure modes. We introduce LongShOTBench, a diagnostic benchmark with open-ended, intent-driven questions; single- and multi-turn dialogues; and tasks requiring multimodal reasoning and agentic tool use across video, audio, and speech. Each item includes a reference answer and graded rubric for interpretable, and traceable evaluation. LongShOTBench is produced via a scalable, human-validated pipeline to ensure coverage and reproducibility. All samples in our LongShOTBench are human-verified and corrected. Furthermore, we present LongShOTAgent, an agentic system that analyzes long videos via preprocessing, search, and iterative refinement. On LongShOTBench, state-of-the-art MLLMs show large gaps: Gemini-2.5-Flash achieves 52.95%, open-source models remain below 30%, and LongShOTAgent attains 44.66%. These results underscore the difficulty of real-world long-form video understanding. LongShOTBench provides a practical, reproducible foundation for evaluating and improving MLLMs. All resources are available on GitHub: https://github.com/mbzuai-oryx/longshot.
title A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.16978