D3: Data Diversity Design for Systematic Generalization in Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Rahimi, Amir, D'Amario, Vanessa, Yamada, Moyuru, Takemoto, Kentaro, Sasaki, Tomotake, Boix, Xavier |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HICO-DET-SG and V-COCO-SG: New Data Splits for Evaluating the Systematic Generalization Performance of Human-Object Interaction Detection Models
by: Takemoto, Kentaro, et al.
Published: (2023)
by: Takemoto, Kentaro, et al.
Published: (2023)
GLoD: Composing Global Contexts and Local Details in Image Generation
by: Yamada, Moyuru
Published: (2024)
by: Yamada, Moyuru
Published: (2024)
Cuestiones de la inclusión educativa. A propósito de la UBV y Misión Sucre
by: Daisy D'Amario
Published: (2009)
by: Daisy D'Amario
Published: (2009)
In-distribution adversarial attacks on object recognition models using gradient-free search
by: Madan, Spandan, et al.
Published: (2021)
by: Madan, Spandan, et al.
Published: (2021)
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)
by: Rosenfeld, Amir, et al.
Published: (2025)
JSynFlow: Japanese Synthesised Flowchart Visual Question Answering Dataset built with Large Language Models
by: Sasaki, Hiroshi
Published: (2026)
by: Sasaki, Hiroshi
Published: (2026)
Exploring Diverse Methods in Visual Question Answering
by: Li, Panfeng, et al.
Published: (2024)
by: Li, Panfeng, et al.
Published: (2024)
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency
by: Dahal, Ashim, et al.
Published: (2025)
by: Dahal, Ashim, et al.
Published: (2025)
Double Machine Learning for Time Series
by: Ciganovic, Milos, et al.
Published: (2026)
by: Ciganovic, Milos, et al.
Published: (2026)
Carbon-Penalised Portfolio Insurance Strategies in a Stochastic Factor Model with Partial Information
by: Colaneri, Katia, et al.
Published: (2025)
by: Colaneri, Katia, et al.
Published: (2025)
STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes
by: Ishihara, Keishi, et al.
Published: (2025)
by: Ishihara, Keishi, et al.
Published: (2025)
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
by: Malik, Sameer, et al.
Published: (2025)
by: Malik, Sameer, et al.
Published: (2025)
Answering Diverse Questions via Text Attached with Key Audio-Visual Clues
by: Ye, Qilang, et al.
Published: (2024)
by: Ye, Qilang, et al.
Published: (2024)
Towards Flexible Evaluation for Generative Visual Question Answering
by: Ji, Huishan, et al.
Published: (2024)
by: Ji, Huishan, et al.
Published: (2024)
Multi-Sourced Compositional Generalization in Visual Question Answering
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
Selectively Answering Visual Questions
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering
by: Das, Deepayan, et al.
Published: (2024)
by: Das, Deepayan, et al.
Published: (2024)
SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering
by: Li, Wenli, et al.
Published: (2026)
by: Li, Wenli, et al.
Published: (2026)
Semantic Graph Consistency: Going Beyond Patches for Regularizing Self-Supervised Vision Transformers
by: Devaguptapu, Chaitanya, et al.
Published: (2024)
by: Devaguptapu, Chaitanya, et al.
Published: (2024)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
by: Su, Tongkun, et al.
Published: (2024)
by: Su, Tongkun, et al.
Published: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
Targeted Visual Prompting for Medical Visual Question Answering
by: Tascon-Morales, Sergio, et al.
Published: (2024)
by: Tascon-Morales, Sergio, et al.
Published: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
by: Fu, Xingyu, et al.
Published: (2023)
by: Fu, Xingyu, et al.
Published: (2023)
Multimodal Rationales for Explainable Visual Question Answering
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Evaluating Variance in Visual Question Answering Benchmarks
by: SR, Nikitha
Published: (2025)
by: SR, Nikitha
Published: (2025)
Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum Learning
by: Zheng, Yuhang, et al.
Published: (2024)
by: Zheng, Yuhang, et al.
Published: (2024)
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
by: Lan, Jian, et al.
Published: (2025)
by: Lan, Jian, et al.
Published: (2025)
RECODE: Reasoning Through Code Generation for Visual Question Answering
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
3D Question Answering for City Scene Understanding
by: Sun, Penglei, et al.
Published: (2024)
by: Sun, Penglei, et al.
Published: (2024)
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
by: You, Wenhao, et al.
Published: (2025)
by: You, Wenhao, et al.
Published: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
Mass and morphology of Emiliania huxleyi coccoliths in the Mediterranean Sea: MedSeA and Meteor M84/3 cruise samples (May 2013, April 2011)
by: D'Amario, Barbara, et al.
Published: (2018)
by: D'Amario, Barbara, et al.
Published: (2018)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
by: Xu, Quanxing, et al.
Published: (2026)
by: Xu, Quanxing, et al.
Published: (2026)
Environmental characteristics and E. huxleyi coccoliths mass and morphology in the Mediterranean Sea during MedSeA and Meteor M84/3 cruises (May 2013, April 2011)
by: D'Amario, Barbara, et al.
Published: (2018)
by: D'Amario, Barbara, et al.
Published: (2018)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
by: Shen, Ruoyue, et al.
Published: (2024)
by: Shen, Ruoyue, et al.
Published: (2024)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
by: Wang, Zining, et al.
Published: (2025)
by: Wang, Zining, et al.
Published: (2025)
Privacy-Aware Document Visual Question Answering
by: Tito, Rubèn, et al.
Published: (2023)
by: Tito, Rubèn, et al.
Published: (2023)
Safe-Construct: Redefining Construction Safety Violation Recognition as 3D Multi-View Engagement Task
by: Chharia, Aviral, et al.
Published: (2025)
by: Chharia, Aviral, et al.
Published: (2025)
Similar Items
-
HICO-DET-SG and V-COCO-SG: New Data Splits for Evaluating the Systematic Generalization Performance of Human-Object Interaction Detection Models
by: Takemoto, Kentaro, et al.
Published: (2023) -
GLoD: Composing Global Contexts and Local Details in Image Generation
by: Yamada, Moyuru
Published: (2024) -
Cuestiones de la inclusión educativa. A propósito de la UBV y Misión Sucre
by: Daisy D'Amario
Published: (2009) -
In-distribution adversarial attacks on object recognition models using gradient-free search
by: Madan, Spandan, et al.
Published: (2021) -
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)