Robust Visual Question Answering: Datasets, Methods, and Future Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Jie, Wang, Pinghui, Kong, Dechen, Wang, Zewei, Liu, Jun, Pei, Hongbin, Zhao, Junzhou |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
by: Ma, Jie, et al.
Published: (2024)
by: Ma, Jie, et al.
Published: (2024)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
by: Castrillón-Santana, Modesto, et al.
Published: (2025)
by: Castrillón-Santana, Modesto, et al.
Published: (2025)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
by: Li, Xiaoyang, et al.
Published: (2025)
by: Li, Xiaoyang, et al.
Published: (2025)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
by: Shah, Nisarg A., et al.
Published: (2025)
by: Shah, Nisarg A., et al.
Published: (2025)
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
by: Han, Yudong, et al.
Published: (2026)
by: Han, Yudong, et al.
Published: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
by: Li, Huibin, et al.
Published: (2025)
by: Li, Huibin, et al.
Published: (2025)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
by: Guo, Yijie, et al.
Published: (2025)
by: Guo, Yijie, et al.
Published: (2025)
Vectra: A New Metric, Dataset, and Model for Visual Quality Assessment in E-Commerce In-Image Machine Translation
by: Wu, Qingyu, et al.
Published: (2026)
by: Wu, Qingyu, et al.
Published: (2026)
A Challenging Benchmark of Anime Style Recognition
by: Li, Haotang, et al.
Published: (2022)
by: Li, Haotang, et al.
Published: (2022)
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
by: Schneider, David, et al.
Published: (2025)
by: Schneider, David, et al.
Published: (2025)
Computer Vision for Clinical Gait Analysis: A Gait Abnormality Video Dataset
by: Ranjan, Rahm, et al.
Published: (2024)
by: Ranjan, Rahm, et al.
Published: (2024)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
by: Fan, Lin, et al.
Published: (2026)
by: Fan, Lin, et al.
Published: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
RAE-NWM: Navigation World Model in Dense Visual Representation Space
by: Zhang, Mingkun, et al.
Published: (2026)
by: Zhang, Mingkun, et al.
Published: (2026)
Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
by: Liu, Jianming, et al.
Published: (2025)
by: Liu, Jianming, et al.
Published: (2025)
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
by: Yang, Xiaoyu, et al.
Published: (2024)
by: Yang, Xiaoyu, et al.
Published: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
A Novel Dataset for Flood Detection Robust to Seasonal Changes in Satellite Imagery
by: Jang, Youngsun, et al.
Published: (2025)
by: Jang, Youngsun, et al.
Published: (2025)
SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
by: Liu, Sicheng, et al.
Published: (2024)
by: Liu, Sicheng, et al.
Published: (2024)
MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
by: Oliveira, Daniel A. P., et al.
Published: (2024)
by: Oliveira, Daniel A. P., et al.
Published: (2024)
Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation
by: Hu, Chenggong, et al.
Published: (2026)
by: Hu, Chenggong, et al.
Published: (2026)
SemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
by: Zheng, Peng, et al.
Published: (2024)
by: Zheng, Peng, et al.
Published: (2024)
Visual Graph Question Answering with ASP and LLMs for Language Parsing
by: Bauer, Jakob Johannes, et al.
Published: (2025)
by: Bauer, Jakob Johannes, et al.
Published: (2025)
GTPBD: A Fine-Grained Global Terraced Parcel and Boundary Dataset
by: Zhang, Zhiwei, et al.
Published: (2025)
by: Zhang, Zhiwei, et al.
Published: (2025)
FocusedAD: Character-centric Movie Audio Description
by: Ye, Xiaojun, et al.
Published: (2025)
by: Ye, Xiaojun, et al.
Published: (2025)
Robust Confidence Intervals in Stereo Matching using Possibility Theory
by: Malinowski, Roman, et al.
Published: (2024)
by: Malinowski, Roman, et al.
Published: (2024)
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
by: Hu, Xuran, et al.
Published: (2026)
by: Hu, Xuran, et al.
Published: (2026)
FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
by: Wang, Zixing, et al.
Published: (2025)
by: Wang, Zixing, et al.
Published: (2025)
Breaking the Resource Wall: Geometry-Guided Sequence Modeling for Efficient Semantic Segmentation
by: Chan, Sheng-Wei, et al.
Published: (2026)
by: Chan, Sheng-Wei, et al.
Published: (2026)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
Lost in Context: The Influence of Context on Feature Attribution Methods for Object Recognition
by: Adhikari, Sayanta, et al.
Published: (2024)
by: Adhikari, Sayanta, et al.
Published: (2024)
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
by: Pandey, Anupam, et al.
Published: (2025)
by: Pandey, Anupam, et al.
Published: (2025)
Pixel-Level Pavement Distress Assessment Using Instance Segmentation
by: Dewick, Logan, et al.
Published: (2026)
by: Dewick, Logan, et al.
Published: (2026)
FoR-Net: Learning to Focus on Hard Regions for Efficient Semantic Segmentation
by: Chan, Sheng-Wei, et al.
Published: (2026)
by: Chan, Sheng-Wei, et al.
Published: (2026)
An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
by: Li, Zhi, et al.
Published: (2025)
by: Li, Zhi, et al.
Published: (2025)
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
by: Chen, Zhangquan, et al.
Published: (2026)
by: Chen, Zhangquan, et al.
Published: (2026)
Similar Items
-
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
by: Ma, Jie, et al.
Published: (2024) -
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
by: Castrillón-Santana, Modesto, et al.
Published: (2025) -
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
by: Li, Xiaoyang, et al.
Published: (2025) -
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
by: Shah, Nisarg A., et al.
Published: (2025) -
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
by: Li, Tao, et al.
Published: (2024)