Admitting Ignorance Helps the Video Question Answering Models to Answer
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Haopeng, Drummond, Tom, Gong, Mingming, Bennamoun, Mohammed, Ke, Qiuhong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
par: Li, Haopeng, et autres
Publié: (2024)
par: Li, Haopeng, et autres
Publié: (2024)
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports
par: Li, Haopeng, et autres
Publié: (2024)
par: Li, Haopeng, et autres
Publié: (2024)
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
par: Liao, Zhaohe, et autres
Publié: (2024)
par: Liao, Zhaohe, et autres
Publié: (2024)
LatentMove: Towards Complex Human Movement Video Generation
par: Taghipour, Ashkan, et autres
Publié: (2025)
par: Taghipour, Ashkan, et autres
Publié: (2025)
TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
par: Liu, Yanan, et autres
Publié: (2025)
par: Liu, Yanan, et autres
Publié: (2025)
Part-aware Unified Representation of Language and Skeleton for Zero-shot Action Recognition
par: Zhu, Anqi, et autres
Publié: (2024)
par: Zhu, Anqi, et autres
Publié: (2024)
Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation
par: Zhu, Jingmin, et autres
Publié: (2025)
par: Zhu, Jingmin, et autres
Publié: (2025)
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
par: Liu, Yexin, et autres
Publié: (2024)
par: Liu, Yexin, et autres
Publié: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
par: Xiao, Junbin, et autres
Publié: (2023)
par: Xiao, Junbin, et autres
Publié: (2023)
An Empirical Study on How Video-LLMs Answer Video Questions
par: Gou, Chenhui, et autres
Publié: (2025)
par: Gou, Chenhui, et autres
Publié: (2025)
Language Model Guided Interpretable Video Action Reasoning
par: Wang, Ning, et autres
Publié: (2024)
par: Wang, Ning, et autres
Publié: (2024)
LongDiff: Training-Free Long Video Generation in One Go
par: Li, Zhuoling, et autres
Publié: (2025)
par: Li, Zhuoling, et autres
Publié: (2025)
DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition
par: Zhu, Jingmin, et autres
Publié: (2025)
par: Zhu, Jingmin, et autres
Publié: (2025)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
par: Miao, Bo, et autres
Publié: (2024)
par: Miao, Bo, et autres
Publié: (2024)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
par: Pintore, Marco, et autres
Publié: (2025)
par: Pintore, Marco, et autres
Publié: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
par: Di, Shangzhe, et autres
Publié: (2025)
par: Di, Shangzhe, et autres
Publié: (2025)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
par: Zhang, Xian, et autres
Publié: (2025)
par: Zhang, Xian, et autres
Publié: (2025)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
par: Song, Enxin, et autres
Publié: (2024)
par: Song, Enxin, et autres
Publié: (2024)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
par: Shabbir, Akashah, et autres
Publié: (2025)
par: Shabbir, Akashah, et autres
Publié: (2025)
Agentic Keyframe Search for Video Question Answering
par: Fan, Sunqi, et autres
Publié: (2025)
par: Fan, Sunqi, et autres
Publié: (2025)
Grounded Question-Answering in Long Egocentric Videos
par: Di, Shangzhe, et autres
Publié: (2023)
par: Di, Shangzhe, et autres
Publié: (2023)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
par: Romero, David, et autres
Publié: (2024)
par: Romero, David, et autres
Publié: (2024)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
par: Nie, Yuxiang, et autres
Publié: (2025)
par: Nie, Yuxiang, et autres
Publié: (2025)
Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
par: Du, Henghui, et autres
Publié: (2025)
par: Du, Henghui, et autres
Publié: (2025)
Narrative Aligned Long Form Video Question Answering
par: Jain, Rahul, et autres
Publié: (2026)
par: Jain, Rahul, et autres
Publié: (2026)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
par: Chaybouti, Sofian, et autres
Publié: (2025)
par: Chaybouti, Sofian, et autres
Publié: (2025)
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
par: Oh, Ju-Young
Publié: (2025)
par: Oh, Ju-Young
Publié: (2025)
ViLA: Efficient Video-Language Alignment for Video Question Answering
par: Wang, Xijun, et autres
Publié: (2023)
par: Wang, Xijun, et autres
Publié: (2023)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
par: Zou, Bo, et autres
Publié: (2024)
par: Zou, Bo, et autres
Publié: (2024)
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
par: Guo, Diandian, et autres
Publié: (2026)
par: Guo, Diandian, et autres
Publié: (2026)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
par: Taghipour, Ashkan, et autres
Publié: (2026)
par: Taghipour, Ashkan, et autres
Publié: (2026)
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
par: Ma, Jie, et autres
Publié: (2024)
par: Ma, Jie, et autres
Publié: (2024)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
par: Xu, Rongtao, et autres
Publié: (2025)
par: Xu, Rongtao, et autres
Publié: (2025)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
par: Fernando, Basura, et autres
Publié: (2025)
par: Fernando, Basura, et autres
Publié: (2025)
Object-centric Video Question Answering with Visual Grounding and Referring
par: Wang, Haochen, et autres
Publié: (2025)
par: Wang, Haochen, et autres
Publié: (2025)
PolySmart @ TRECVid 2024 Medical Video Question Answering
par: Wu, Jiaxin, et autres
Publié: (2024)
par: Wu, Jiaxin, et autres
Publié: (2024)
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
par: Kugo, Noriyuki, et autres
Publié: (2024)
par: Kugo, Noriyuki, et autres
Publié: (2024)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
par: Shi, Mengqi, et autres
Publié: (2026)
par: Shi, Mengqi, et autres
Publié: (2026)
HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models
par: Bai, Xiangyu, et autres
Publié: (2026)
par: Bai, Xiangyu, et autres
Publié: (2026)
CinePile: A Long Video Question Answering Dataset and Benchmark
par: Rawal, Ruchit, et autres
Publié: (2024)
par: Rawal, Ruchit, et autres
Publié: (2024)
Documents similaires
-
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
par: Li, Haopeng, et autres
Publié: (2024) -
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports
par: Li, Haopeng, et autres
Publié: (2024) -
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
par: Liao, Zhaohe, et autres
Publié: (2024) -
LatentMove: Towards Complex Human Movement Video Generation
par: Taghipour, Ashkan, et autres
Publié: (2025) -
TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
par: Liu, Yanan, et autres
Publié: (2025)