Acknowledging Focus Ambiguity in Visual Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Chongyan, Tseng, Yu-Yun, Li, Zhuoheng, Venkatesh, Anush, Gurari, Danna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fully Authentic Visual Question Answering Dataset from Online Communities
by: Chen, Chongyan, et al.
Published: (2023)
by: Chen, Chongyan, et al.
Published: (2023)
An Evaluation of GPT-4V and Gemini in Online VQA
by: Liu, Mengchen, et al.
Published: (2023)
by: Liu, Mengchen, et al.
Published: (2023)
Long-Form Answers to Visual Questions from Blind and Low Vision People
by: Huh, Mina, et al.
Published: (2024)
by: Huh, Mina, et al.
Published: (2024)
BIV-Priv-Seg: Locating Private Content in Images Taken by People With Visual Impairments
by: Tseng, Yu-Yun, et al.
Published: (2024)
by: Tseng, Yu-Yun, et al.
Published: (2024)
Collecting Consistently High Quality Object Tracks with Minimal Human Involvement by Using Self-Supervised Learning to Detect Tracker Errors
by: Anjum, Samreen, et al.
Published: (2024)
by: Anjum, Samreen, et al.
Published: (2024)
PartStickers: Generating Parts of Objects for Rapid Prototyping
by: Zhou, Mo, et al.
Published: (2025)
by: Zhou, Mo, et al.
Published: (2025)
SPIN: Hierarchical Segmentation with Subpart Granularity in Natural Images
by: Myers-Dean, Josh, et al.
Published: (2024)
by: Myers-Dean, Josh, et al.
Published: (2024)
Interpreting COVID Lateral Flow Tests' Results with Foundation Models
by: Pandey, Stuti, et al.
Published: (2024)
by: Pandey, Stuti, et al.
Published: (2024)
Logit-Based Losses Limit the Effectiveness of Feature Knowledge Distillation
by: Cooper, Nicholas, et al.
Published: (2025)
by: Cooper, Nicholas, et al.
Published: (2025)
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
by: Jian, Pu, et al.
Published: (2025)
by: Jian, Pu, et al.
Published: (2025)
Hierarchical Instance Tracking to Balance Privacy Preservation with Accessible Information
by: Prasad, Neelima, et al.
Published: (2025)
by: Prasad, Neelima, et al.
Published: (2025)
Extended to Reality: Prompt Injection in 3D Environments
by: Li, Zhuoheng, et al.
Published: (2026)
by: Li, Zhuoheng, et al.
Published: (2026)
Improving Few-Shot Change Detection Visual Question Answering via Decision-Ambiguity-guided Reinforcement Fine-Tuning
by: Dong, Fuyu, et al.
Published: (2025)
by: Dong, Fuyu, et al.
Published: (2025)
ConFoThinking: Consolidated Focused Attention Driven Thinking for Visual Question Answering
by: Wu, Zhaodong, et al.
Published: (2026)
by: Wu, Zhaodong, et al.
Published: (2026)
DARK: Denoising, Amplification, Restoration Kit
by: Li, Zhuoheng, et al.
Published: (2024)
by: Li, Zhuoheng, et al.
Published: (2024)
Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
by: Bai, Ziyi, et al.
Published: (2024)
by: Bai, Ziyi, et al.
Published: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
FVAR: Visual Autoregressive Modeling via Next Focus Prediction
by: Li, Xiaofan, et al.
Published: (2025)
by: Li, Xiaofan, et al.
Published: (2025)
Consistency and Uncertainty: Identifying Unreliable Responses From Black-Box Vision-Language Models for Selective Visual Question Answering
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
by: Xue, Junxiao, et al.
Published: (2024)
by: Xue, Junxiao, et al.
Published: (2024)
Eliminating Feature Ambiguity for Few-Shot Segmentation
by: Xu, Qianxiong, et al.
Published: (2024)
by: Xu, Qianxiong, et al.
Published: (2024)
Focus on Local: Finding Reliable Discriminative Regions for Visual Place Recognition
by: Wang, Changwei, et al.
Published: (2025)
by: Wang, Changwei, et al.
Published: (2025)
Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
DriveLM: Driving with Graph Visual Question Answering
by: Sima, Chonghao, et al.
Published: (2023)
by: Sima, Chonghao, et al.
Published: (2023)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
Multimodal Rationales for Explainable Visual Question Answering
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
READ-Net: Clarifying Emotional Ambiguity via Adaptive Feature Recalibration for Audio-Visual Depression Detection
by: Chen, Chenglizhao, et al.
Published: (2026)
by: Chen, Chenglizhao, et al.
Published: (2026)
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
by: Li, Yuyi, et al.
Published: (2025)
by: Li, Yuyi, et al.
Published: (2025)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
Revisiting Photometric Ambiguity for Accurate Gaussian-Splatting Surface Reconstruction
by: Li, Jiahe, et al.
Published: (2026)
by: Li, Jiahe, et al.
Published: (2026)
VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering
by: Chen, Jiayi, et al.
Published: (2026)
by: Chen, Jiayi, et al.
Published: (2026)
A Cat Is A Cat (Not A Dog!): Unraveling Information Mix-ups in Text-to-Image Encoders through Causal Analysis and Embedding Optimization
by: Chen, Chieh-Yun, et al.
Published: (2024)
by: Chen, Chieh-Yun, et al.
Published: (2024)
TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
by: Kim, Jongha, et al.
Published: (2025)
by: Kim, Jongha, et al.
Published: (2025)
Targeted Visual Prompting for Medical Visual Question Answering
by: Tascon-Morales, Sergio, et al.
Published: (2024)
by: Tascon-Morales, Sergio, et al.
Published: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
Knowledge-aware Visual Question Generation for Remote Sensing Images
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
HyperNVD: Accelerating Neural Video Decomposition via Hypernetworks
by: Pilligua, Maria, et al.
Published: (2025)
by: Pilligua, Maria, et al.
Published: (2025)
Similar Items
-
Fully Authentic Visual Question Answering Dataset from Online Communities
by: Chen, Chongyan, et al.
Published: (2023) -
An Evaluation of GPT-4V and Gemini in Online VQA
by: Liu, Mengchen, et al.
Published: (2023) -
Long-Form Answers to Visual Questions from Blind and Low Vision People
by: Huh, Mina, et al.
Published: (2024) -
BIV-Priv-Seg: Locating Private Content in Images Taken by People With Visual Impairments
by: Tseng, Yu-Yun, et al.
Published: (2024) -
Collecting Consistently High Quality Object Tracks with Minimal Human Involvement by Using Self-Supervised Learning to Detect Tracker Errors
by: Anjum, Samreen, et al.
Published: (2024)