VisionTrap: Unanswerable Questions On Visual Data
Fuente:
arXiv
Saved in:
| Main Authors: | Saadat, Asir, Aziz, Syem, Mahmud, Shahriar, Mahi, Abdullah Ibne Masud, Ahmed, Sabbir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions
by: Moon, Seokha, et al.
Published: (2024)
by: Moon, Seokha, et al.
Published: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems
by: Saadat, Asir, et al.
Published: (2024)
by: Saadat, Asir, et al.
Published: (2024)
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
by: He, Xingwei, et al.
Published: (2024)
by: He, Xingwei, et al.
Published: (2024)
AQUA20: A Benchmark Dataset for Underwater Species Classification under Challenging Conditions
by: Fuad, Taufikur Rahman, et al.
Published: (2025)
by: Fuad, Taufikur Rahman, et al.
Published: (2025)
CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering
by: Vardi, Ben, et al.
Published: (2025)
by: Vardi, Ben, et al.
Published: (2025)
Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents
by: Napolitano, Davide, et al.
Published: (2025)
by: Napolitano, Davide, et al.
Published: (2025)
MangoLeafViT: Leveraging Lightweight Vision Transformer with Runtime Augmentation for Efficient Mango Leaf Disease Classification
by: Chowdhury, Rafi Hassan, et al.
Published: (2025)
by: Chowdhury, Rafi Hassan, et al.
Published: (2025)
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
by: Zhu, Yanxu, et al.
Published: (2025)
by: Zhu, Yanxu, et al.
Published: (2025)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
by: Rahman, Md Ashikur, et al.
Published: (2026)
by: Rahman, Md Ashikur, et al.
Published: (2026)
Visual Bias and Interpretability in Deep Learning for Dermatological Image Analysis
by: Taufik, Enam Ahmed, et al.
Published: (2025)
by: Taufik, Enam Ahmed, et al.
Published: (2025)
Automatic Vehicle Detection using DETR: A Transformer-Based Approach for Navigating Treacherous Roads
by: Fahad, Istiaq Ahmed, et al.
Published: (2025)
by: Fahad, Istiaq Ahmed, et al.
Published: (2025)
Computational Framework for Estimating Relative Gaussian Blur Kernels between Image Pairs
by: Saadat, Akbar
Published: (2026)
by: Saadat, Akbar
Published: (2026)
Defocus Aberration Theory Confirms Gaussian Model in Most Imaging Devices
by: Saadat, Akbar
Published: (2026)
by: Saadat, Akbar
Published: (2026)
DExNet: Combining Observations of Domain Adapted Critics for Leaf Disease Classification with Limited Data
by: Ahmed, Sabbir, et al.
Published: (2025)
by: Ahmed, Sabbir, et al.
Published: (2025)
Beyond Dominant Patches: Spatial Credit Redistribution For Grounded Vision-Language Models
by: Samin, Niamul Hassan, et al.
Published: (2026)
by: Samin, Niamul Hassan, et al.
Published: (2026)
EVCC: Enhanced Vision Transformer-ConvNeXt-CoAtNet Fusion for Classification
by: Hasan, Kazi Reyazul, et al.
Published: (2025)
by: Hasan, Kazi Reyazul, et al.
Published: (2025)
Advanced Vision Transformers and Open-Set Learning for Robust Mosquito Classification: A Novel Approach to Entomological Studies
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
by: Karim, Ahmed Akib Jawad, et al.
Published: (2024)
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
by: Shah, Monika, et al.
Published: (2025)
by: Shah, Monika, et al.
Published: (2025)
Skin Cancer Segmentation and Classification Using Vision Transformer for Automatic Analysis in Dermatoscopy-based Non-invasive Digital System
by: Himel, Galib Muhammad Shahriar, et al.
Published: (2024)
by: Himel, Galib Muhammad Shahriar, et al.
Published: (2024)
OncoVision: Integrating Mammography and Clinical Data through Attention-Driven Multimodal AI for Enhanced Breast Cancer Diagnosis
by: Ahmed, Istiak, et al.
Published: (2025)
by: Ahmed, Istiak, et al.
Published: (2025)
Are MLMs Trapped in the Visual Room?
by: Zhang, Yazhou, et al.
Published: (2025)
by: Zhang, Yazhou, et al.
Published: (2025)
TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection
by: Abdullah, Ahmed, et al.
Published: (2026)
by: Abdullah, Ahmed, et al.
Published: (2026)
Unified Alignment Protocol: Making Sense of the Unlabeled Data in New Domains
by: Ahmed, Sabbir, et al.
Published: (2025)
by: Ahmed, Sabbir, et al.
Published: (2025)
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
by: Jian, Pu, et al.
Published: (2025)
by: Jian, Pu, et al.
Published: (2025)
Large Vision-Language Models for Remote Sensing Visual Question Answering
by: Siripong, Surasakdi, et al.
Published: (2024)
by: Siripong, Surasakdi, et al.
Published: (2024)
Activator: GLU Activation Function as the Core Component of a Vision Transformer
by: Abdullah, Abdullah Nazhat, et al.
Published: (2024)
by: Abdullah, Abdullah Nazhat, et al.
Published: (2024)
BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users
by: Cheng, Wanyin, et al.
Published: (2025)
by: Cheng, Wanyin, et al.
Published: (2025)
Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
by: Kapuriya, Janak, et al.
Published: (2025)
by: Kapuriya, Janak, et al.
Published: (2025)
Few-Shot Image Classification and Segmentation as Visual Question Answering Using Vision-Language Models
by: Meng, Tian, et al.
Published: (2024)
by: Meng, Tian, et al.
Published: (2024)
Fusion of Domain-Adapted Vision and Language Models for Medical Visual Question Answering
by: Ha, Cuong Nhat, et al.
Published: (2024)
by: Ha, Cuong Nhat, et al.
Published: (2024)
Long-Form Answers to Visual Questions from Blind and Low Vision People
by: Huh, Mina, et al.
Published: (2024)
by: Huh, Mina, et al.
Published: (2024)
Characterizing Disparity Between Edge Models and High-Accuracy Base Models for Vision Tasks
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
by: Kim, Younggun, et al.
Published: (2025)
by: Kim, Younggun, et al.
Published: (2025)
VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models
by: Wang, Chenyu, et al.
Published: (2026)
by: Wang, Chenyu, et al.
Published: (2026)
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
by: Lan, Jian, et al.
Published: (2025)
by: Lan, Jian, et al.
Published: (2025)
Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum Learning
by: Zheng, Yuhang, et al.
Published: (2024)
by: Zheng, Yuhang, et al.
Published: (2024)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
PCFEx: Point Cloud Feature Extraction for Graph Neural Networks
by: Masud, Abdullah Al, et al.
Published: (2026)
by: Masud, Abdullah Al, et al.
Published: (2026)
Similar Items
-
VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions
by: Moon, Seokha, et al.
Published: (2024) -
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024) -
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems
by: Saadat, Asir, et al.
Published: (2024) -
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024) -
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
by: He, Xingwei, et al.
Published: (2024)