Saved in:
| Main Authors: | Jain, Gautam Kumar, Markgraf, Carsten, Stähler, Julian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.22560 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
by: Madaka, Madhuri Latha, et al.
Published: (2025)
by: Madaka, Madhuri Latha, et al.
Published: (2025)
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
by: Etchegaray, Djamahl, et al.
Published: (2025)
by: Etchegaray, Djamahl, et al.
Published: (2025)
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
by: Chen, Yixiong, et al.
Published: (2026)
by: Chen, Yixiong, et al.
Published: (2026)
Kvasir-VQA: A Text-Image Pair GI Tract Dataset
by: Gautam, Sushant, et al.
Published: (2024)
by: Gautam, Sushant, et al.
Published: (2024)
Understanding the Role of the Projector in Knowledge Distillation
by: Miles, Roy, et al.
Published: (2023)
by: Miles, Roy, et al.
Published: (2023)
A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
Explicit Correlation Learning for Generalizable Cross-Modal Deepfake Detection
by: Yu, Cai, et al.
Published: (2024)
by: Yu, Cai, et al.
Published: (2024)
Context-Gated Cross-Modal Perception with Visual Mamba for PET-CT Lung Tumor Segmentation
by: Ayllón, Elena Mulero, et al.
Published: (2025)
by: Ayllón, Elena Mulero, et al.
Published: (2025)
HELM: Hierarchical and Explicit Label Modeling with Graph Learning for Multi-Label Image Classification
by: Stoimchev, Marjan, et al.
Published: (2026)
by: Stoimchev, Marjan, et al.
Published: (2026)
Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling
by: Zou, Hongjian, et al.
Published: (2026)
by: Zou, Hongjian, et al.
Published: (2026)
Improving VQA Reliability: A Dual-Assessment Approach with Self-Reflection and Cross-Model Verification
by: Wu, Xixian, et al.
Published: (2025)
by: Wu, Xixian, et al.
Published: (2025)
Training-Free Representation Guidance for Diffusion Models with a Representation Alignment Projector
by: Zu, Wenqiang, et al.
Published: (2026)
by: Zu, Wenqiang, et al.
Published: (2026)
Device-aware Optical Adversarial Attack for a Portable Projector-camera System
by: Jiang, Ning, et al.
Published: (2025)
by: Jiang, Ning, et al.
Published: (2025)
QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
by: Li, Zhongyang, et al.
Published: (2026)
by: Li, Zhongyang, et al.
Published: (2026)
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA
by: Pandey, Ayush, et al.
Published: (2025)
by: Pandey, Ayush, et al.
Published: (2025)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
by: Sun, Boyuan, et al.
Published: (2025)
by: Sun, Boyuan, et al.
Published: (2025)
Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
by: Yang, Xiaobo, et al.
Published: (2025)
by: Yang, Xiaobo, et al.
Published: (2025)
Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru
by: Cusipuma, Dunant, et al.
Published: (2025)
by: Cusipuma, Dunant, et al.
Published: (2025)
Is ChatGPT-5 Ready for Mammogram VQA?
by: Li, Qiang, et al.
Published: (2025)
by: Li, Qiang, et al.
Published: (2025)
Advancing Surgical VQA with Scene Graph Knowledge
by: Yuan, Kun, et al.
Published: (2023)
by: Yuan, Kun, et al.
Published: (2023)
Hyperspectral Imaging-Based Perception in Autonomous Driving Scenarios: Benchmarking Baseline Semantic Segmentation Models
by: Shah, Imad Ali, et al.
Published: (2024)
by: Shah, Imad Ali, et al.
Published: (2024)
DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving
by: Fu, Yongjie, et al.
Published: (2024)
by: Fu, Yongjie, et al.
Published: (2024)
Dual Causal Inference: Integrating Backdoor Adjustment and Instrumental Variable Learning for Medical VQA
by: Xu, Zibo, et al.
Published: (2026)
by: Xu, Zibo, et al.
Published: (2026)
Honeybee: Locality-enhanced Projector for Multimodal LLM
by: Cha, Junbum, et al.
Published: (2023)
by: Cha, Junbum, et al.
Published: (2023)
An Evaluation of GPT-4V and Gemini in Online VQA
by: Liu, Mengchen, et al.
Published: (2023)
by: Liu, Mengchen, et al.
Published: (2023)
KNVQA: A Benchmark for evaluation knowledge-based VQA
by: Cheng, Sirui, et al.
Published: (2023)
by: Cheng, Sirui, et al.
Published: (2023)
AID4AD: Aerial Image Data for Automated Driving Perception
by: Lengerer, Daniel, et al.
Published: (2025)
by: Lengerer, Daniel, et al.
Published: (2025)
SkeleGuide: Explicit Skeleton Reasoning for Context-Aware Human-in-Place Image Synthesis
by: Wu, Chuqiao, et al.
Published: (2026)
by: Wu, Chuqiao, et al.
Published: (2026)
Back to the Baseline: Examining Baseline Effects on Explainability Metrics
by: Picard, Agustin Martin, et al.
Published: (2025)
by: Picard, Agustin Martin, et al.
Published: (2025)
ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
R^3-VQA: "Read the Room" by Video Social Reasoning
by: Niu, Lixing, et al.
Published: (2025)
by: Niu, Lixing, et al.
Published: (2025)
MISS: A Generative Pretraining and Finetuning Approach for Med-VQA
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Better Supervised Fine-tuning for VQA: Integer-Only Loss
by: Qian, Baihong, et al.
Published: (2025)
by: Qian, Baihong, et al.
Published: (2025)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
Fast globally optimal Truncated Least Squares point cloud registration with fixed rotation axis
by: Ivanov, Ivo, et al.
Published: (2025)
by: Ivanov, Ivo, et al.
Published: (2025)
A Framework for Cross-Domain Generalization in Coronary Artery Calcium Scoring Across Gated and Non-Gated Computed Tomography
by: Gokmen, Mahmut S., et al.
Published: (2026)
by: Gokmen, Mahmut S., et al.
Published: (2026)
Pedestrian Attribute Recognition via Hierarchical Cross-Modality HyperGraph Learning
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
A Data-Centric Vision Transformer Baseline for SAR Sea Ice Classification
by: Mike-Ewewie, David, et al.
Published: (2026)
by: Mike-Ewewie, David, et al.
Published: (2026)
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
by: Srivastava, Archita, et al.
Published: (2025)
by: Srivastava, Archita, et al.
Published: (2025)
Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction
by: Lengerer, Daniel, et al.
Published: (2026)
by: Lengerer, Daniel, et al.
Published: (2026)
Similar Items
-
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
by: Madaka, Madhuri Latha, et al.
Published: (2025) -
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
by: Etchegaray, Djamahl, et al.
Published: (2025) -
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
by: Chen, Yixiong, et al.
Published: (2026) -
Kvasir-VQA: A Text-Image Pair GI Tract Dataset
by: Gautam, Sushant, et al.
Published: (2024) -
Understanding the Role of the Projector in Knowledge Distillation
by: Miles, Roy, et al.
Published: (2023)