Similar Items
Text-guided Fine-Grained Video Anomaly Understanding
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
On the Role of Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2024)
by: Reich, Daniel, et al.
Published: (2024)
Enhancing Document VQA Models via Retrieval-Augmented Generation
by: López, Eric, et al.
Published: (2025)
by: López, Eric, et al.
Published: (2025)
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
by: Dong, Jihao, et al.
Published: (2024)
by: Dong, Jihao, et al.
Published: (2024)
MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment
by: wang, Shuo, et al.
Published: (2025)
by: wang, Shuo, et al.
Published: (2025)
MDS-VQA: Model-Informed Data Selection for Video Quality Assessment
by: Zou, Jian, et al.
Published: (2026)
by: Zou, Jian, et al.
Published: (2026)
Understanding Multi-Agent Reasoning with Large Language Models for Cartoon VQA
by: Wu, Tong, et al.
Published: (2026)
by: Wu, Tong, et al.
Published: (2026)
Modularized Zero-shot VQA with Pre-trained Models
by: Cao, Rui, et al.
Published: (2023)
by: Cao, Rui, et al.
Published: (2023)
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving
by: Liu, Yibo, et al.
Published: (2024)
by: Liu, Yibo, et al.
Published: (2024)
Exploring OCR-augmented Generation for Bilingual VQA
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Knowledge Condensation and Reasoning for Knowledge-based VQA
by: Hao, Dongze, et al.
Published: (2024)
by: Hao, Dongze, et al.
Published: (2024)
Measuring Faithful and Plausible Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2023)
by: Reich, Daniel, et al.
Published: (2023)
DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labels
by: Guo, Erjian, et al.
Published: (2025)
by: Guo, Erjian, et al.
Published: (2025)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
by: Kahl, Kim-Celine, et al.
Published: (2024)
by: Kahl, Kim-Celine, et al.
Published: (2024)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
SurgCheck: Do Vision-Language Models Really Look at Images in Surgical VQA?
by: Shin, Jongmin, et al.
Published: (2026)
by: Shin, Jongmin, et al.
Published: (2026)
HLTCOE Evaluation Team at TREC 2025: VQA Track
by: Zhang, Dengjia, et al.
Published: (2025)
by: Zhang, Dengjia, et al.
Published: (2025)
SplatTalk: 3D VQA with Gaussian Splatting
by: Thai, Anh, et al.
Published: (2025)
by: Thai, Anh, et al.
Published: (2025)
OpenView: Empowering MLLMs with Out-of-view VQA
by: Chen, Qixiang, et al.
Published: (2025)
by: Chen, Qixiang, et al.
Published: (2025)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
Uncovering the Full Potential of Visual Grounding Methods in VQA
by: Reich, Daniel, et al.
Published: (2024)
by: Reich, Daniel, et al.
Published: (2024)
MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
MA-Bench: Towards Fine-grained Micro-Action Understanding
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
Embodied Scene Understanding for Vision Language Models via MetaVQA
by: Wang, Weizhen, et al.
Published: (2025)
by: Wang, Weizhen, et al.
Published: (2025)
CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering
by: Hong, Yuyang, et al.
Published: (2026)
by: Hong, Yuyang, et al.
Published: (2026)
Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
by: Zhou, Xunchu, et al.
Published: (2024)
by: Zhou, Xunchu, et al.
Published: (2024)
UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models
by: Guo, Yangyang, et al.
Published: (2023)
by: Guo, Yangyang, et al.
Published: (2023)
Learn 3D VQA Better with Active Selection and Reannotation
by: Zhou, Shengli, et al.
Published: (2025)
by: Zhou, Shengli, et al.
Published: (2025)
Integrating Query-aware Segmentation and Cross-Attention for Robust VQA
by: Choi, Wonjun, et al.
Published: (2024)
by: Choi, Wonjun, et al.
Published: (2024)
Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
by: Miao, Qiaomu, et al.
Published: (2024)
by: Miao, Qiaomu, et al.
Published: (2024)
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
Enhancing Vision-Language Model with Unmasked Token Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
by: Shahgir, Haz Sameen, et al.
Published: (2024)
by: Shahgir, Haz Sameen, et al.
Published: (2024)
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models
by: Sayem, MD Khalequzzaman Chowdhury, et al.
Published: (2026)
by: Sayem, MD Khalequzzaman Chowdhury, et al.
Published: (2026)
Similar Items
-
Text-guided Fine-Grained Video Anomaly Understanding
by: Gu, Jihao, et al.
Published: (2025) -
On the Role of Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2024) -
Enhancing Document VQA Models via Retrieval-Augmented Generation
by: López, Eric, et al.
Published: (2025) -
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
by: Dong, Jihao, et al.
Published: (2024) -
MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment
by: wang, Shuo, et al.
Published: (2025)