Saved in:
| Main Authors: | Xu, Yinsong, Jing, Wei, Zhang, Liuxin, Lv, Wanjun, Li, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.29402 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
by: Yang, Sicheng, et al.
Published: (2026)
by: Yang, Sicheng, et al.
Published: (2026)
R^3-VQA: "Read the Room" by Video Social Reasoning
by: Niu, Lixing, et al.
Published: (2025)
by: Niu, Lixing, et al.
Published: (2025)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge
by: Taluzzi, Agnese, et al.
Published: (2025)
by: Taluzzi, Agnese, et al.
Published: (2025)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
by: Du, Wanjun, et al.
Published: (2026)
by: Du, Wanjun, et al.
Published: (2026)
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space
by: Singh, Anshul, et al.
Published: (2025)
by: Singh, Anshul, et al.
Published: (2025)
EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge
by: Chen, Zhiwei, et al.
Published: (2026)
by: Chen, Zhiwei, et al.
Published: (2026)
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
by: Lv, Guannan, et al.
Published: (2026)
by: Lv, Guannan, et al.
Published: (2026)
From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
by: Ishmam, Md Farhan, et al.
Published: (2023)
by: Ishmam, Md Farhan, et al.
Published: (2023)
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models
by: Ge, Qihang, et al.
Published: (2024)
by: Ge, Qihang, et al.
Published: (2024)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
by: Ma, Jiatong, et al.
Published: (2026)
by: Ma, Jiatong, et al.
Published: (2026)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
by: Wang, Xingyue, et al.
Published: (2026)
by: Wang, Xingyue, et al.
Published: (2026)
CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
by: Han, Hongyong, et al.
Published: (2025)
by: Han, Hongyong, et al.
Published: (2025)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
by: Vu, Sinh Trong, et al.
Published: (2025)
by: Vu, Sinh Trong, et al.
Published: (2025)
Geometry-Aware Semantic Reasoning for Training Free Video Anomaly Detection
by: Zia, Ali, et al.
Published: (2026)
by: Zia, Ali, et al.
Published: (2026)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024)
by: Chen, Pingyi, et al.
Published: (2024)
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA
by: Carlini, Luca, et al.
Published: (2025)
by: Carlini, Luca, et al.
Published: (2025)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)
by: Qin, Jialong, et al.
Published: (2025)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
by: Cho, Janghoon, et al.
Published: (2025)
by: Cho, Janghoon, et al.
Published: (2025)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025)
by: Ding, Yang, et al.
Published: (2025)
Knowledge Generation for Zero-shot Knowledge-based VQA
by: Cao, Rui, et al.
Published: (2024)
by: Cao, Rui, et al.
Published: (2024)
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
by: Chen, Xupeng, et al.
Published: (2024)
by: Chen, Xupeng, et al.
Published: (2024)
TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
by: Kim, Yoonsik, et al.
Published: (2024)
by: Kim, Yoonsik, et al.
Published: (2024)
MaS-VQA: A Mask-and-Select Framework for Knowledge-Based Visual Question Answering
by: Mao, Xianwei, et al.
Published: (2026)
by: Mao, Xianwei, et al.
Published: (2026)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
by: Wang, Xucheng, et al.
Published: (2026)
by: Wang, Xucheng, et al.
Published: (2026)
HD-EPIC: A Highly-Detailed Egocentric Video Dataset
by: Perrett, Toby, et al.
Published: (2025)
by: Perrett, Toby, et al.
Published: (2025)
ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
by: Zhou, Yuan, et al.
Published: (2025)
by: Zhou, Yuan, et al.
Published: (2025)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
by: Luo, Yongdong, et al.
Published: (2024)
by: Luo, Yongdong, et al.
Published: (2024)
VISD: Enhancing Video Reasoning via Structured Self-Distillation
by: Lin, Hao, et al.
Published: (2026)
by: Lin, Hao, et al.
Published: (2026)
Semantic Event Graphs for Long-Form Video Question Answering
by: Dixit, Aradhya, et al.
Published: (2026)
by: Dixit, Aradhya, et al.
Published: (2026)
Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding
by: Guo, Weiyu, et al.
Published: (2025)
by: Guo, Weiyu, et al.
Published: (2025)
BioD2C: A Dual-level Semantic Consistency Constraint Framework for Biomedical VQA
by: Ji, Zhengyang, et al.
Published: (2025)
by: Ji, Zhengyang, et al.
Published: (2025)
Solution for CVPR 2024 UG2+ Challenge Track on All Weather Semantic Segmentation
by: Yu, Jun, et al.
Published: (2024)
by: Yu, Jun, et al.
Published: (2024)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
by: Han, Kai, et al.
Published: (2024)
by: Han, Kai, et al.
Published: (2024)
Similar Items
-
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
by: Yang, Sicheng, et al.
Published: (2026) -
R^3-VQA: "Read the Room" by Video Social Reasoning
by: Niu, Lixing, et al.
Published: (2025) -
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024) -
From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge
by: Taluzzi, Agnese, et al.
Published: (2025) -
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)