TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akl, Ahmed, Khamis, Abdelwahed, Wang, Zhe, Cheraghian, Ali, Khalifa, Sara, Wang, Kewen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HIME: Mitigating Object Hallucinations in LVLMs via Hallucination Insensitivity Model Editing
von: Akl, Ahmed, et al.
Veröffentlicht: (2026)
von: Akl, Ahmed, et al.
Veröffentlicht: (2026)
SteerSeg: Attention Steering for Reasoning Video Segmentation
von: Cheraghian, Ali, et al.
Veröffentlicht: (2026)
von: Cheraghian, Ali, et al.
Veröffentlicht: (2026)
Test-Time Adaptation for Anomaly Segmentation via Topology-Aware Optimal Transport Chaining
von: Zia, Ali, et al.
Veröffentlicht: (2026)
von: Zia, Ali, et al.
Veröffentlicht: (2026)
Questioning the Stability of Visual Question Answering
von: Rosenfeld, Amir, et al.
Veröffentlicht: (2025)
von: Rosenfeld, Amir, et al.
Veröffentlicht: (2025)
NeuralPrefix: A Zero-shot Sensory Data Imputation Plugin
von: Khamis, Abdelwahed, et al.
Veröffentlicht: (2025)
von: Khamis, Abdelwahed, et al.
Veröffentlicht: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
von: Wang, Yuduo, et al.
Veröffentlicht: (2023)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
von: Souibgui, Mohamed Ali, et al.
Veröffentlicht: (2025)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
von: Indrehus, Kjetil, et al.
Veröffentlicht: (2026)
von: Indrehus, Kjetil, et al.
Veröffentlicht: (2026)
BERT-VQA: Visual Question Answering on Plots
von: Vu, Tai, et al.
Veröffentlicht: (2025)
von: Vu, Tai, et al.
Veröffentlicht: (2025)
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering
von: Weng, Weixi, et al.
Veröffentlicht: (2024)
von: Weng, Weixi, et al.
Veröffentlicht: (2024)
Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum Learning
von: Zheng, Yuhang, et al.
Veröffentlicht: (2024)
von: Zheng, Yuhang, et al.
Veröffentlicht: (2024)
Describe Anything Model for Visual Question Answering on Text-rich Images
von: Vu, Yen-Linh, et al.
Veröffentlicht: (2025)
von: Vu, Yen-Linh, et al.
Veröffentlicht: (2025)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
Privacy-Aware Document Visual Question Answering
von: Tito, Rubèn, et al.
Veröffentlicht: (2023)
von: Tito, Rubèn, et al.
Veröffentlicht: (2023)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
von: Chintapatla, Ishant, et al.
Veröffentlicht: (2025)
von: Chintapatla, Ishant, et al.
Veröffentlicht: (2025)
Expert Knowledge-Aware Image Difference Graph Representation Learning for Difference-Aware Medical Visual Question Answering
von: Hu, Xinyue, et al.
Veröffentlicht: (2023)
von: Hu, Xinyue, et al.
Veröffentlicht: (2023)
Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering
von: Alavi, Ali
Veröffentlicht: (2026)
von: Alavi, Ali
Veröffentlicht: (2026)
Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
von: Hartsock, Iryna, et al.
Veröffentlicht: (2024)
von: Hartsock, Iryna, et al.
Veröffentlicht: (2024)
TLG: Temporal-Logic Grounding for Video Question Answering via Source-Annotation Reconstruction and Category-Targeted Reasoning
von: Alavi, Ali
Veröffentlicht: (2026)
von: Alavi, Ali
Veröffentlicht: (2026)
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
von: Li, Bingxin
Veröffentlicht: (2025)
von: Li, Bingxin
Veröffentlicht: (2025)
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
von: Mamaghan, Amir Mohammad Karimi, et al.
Veröffentlicht: (2024)
von: Mamaghan, Amir Mohammad Karimi, et al.
Veröffentlicht: (2024)
RECODE: Reasoning Through Code Generation for Visual Question Answering
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
von: Van-Dinh, Tue-Thu, et al.
Veröffentlicht: (2025)
von: Van-Dinh, Tue-Thu, et al.
Veröffentlicht: (2025)
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection
von: Lei, Zhanhe, et al.
Veröffentlicht: (2026)
von: Lei, Zhanhe, et al.
Veröffentlicht: (2026)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
Exploring Diverse Methods in Visual Question Answering
von: Li, Panfeng, et al.
Veröffentlicht: (2024)
von: Li, Panfeng, et al.
Veröffentlicht: (2024)
STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
von: Li, Xu, et al.
Veröffentlicht: (2025)
von: Li, Xu, et al.
Veröffentlicht: (2025)
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
von: Huai, Tianyu, et al.
Veröffentlicht: (2025)
von: Huai, Tianyu, et al.
Veröffentlicht: (2025)
Tab2Visual: Overcoming Limited Data in Tabular Data Classification Using Deep Learning with Visual Representations
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2025)
von: Mamdouh, Ahmed, et al.
Veröffentlicht: (2025)
CLoE: Curriculum Learning on Endoscopic Images for Robust MES Classification
von: Ozdemir, Zeynep, et al.
Veröffentlicht: (2025)
von: Ozdemir, Zeynep, et al.
Veröffentlicht: (2025)
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
von: Diao, Xingjian, et al.
Veröffentlicht: (2025)
Improving Video Question Answering through query-based frame selection
von: Patil, Himanshu, et al.
Veröffentlicht: (2026)
von: Patil, Himanshu, et al.
Veröffentlicht: (2026)
Enhancing Multi-Image Question Answering via Submodular Subset Selection
von: Sharma, Aaryan, et al.
Veröffentlicht: (2025)
von: Sharma, Aaryan, et al.
Veröffentlicht: (2025)
ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
von: Chahe, Amirhosein, et al.
Veröffentlicht: (2025)
von: Chahe, Amirhosein, et al.
Veröffentlicht: (2025)
Multi-Task Learning for Integrated Automated Contouring and Voxel-Based Dose Prediction in Radiotherapy
von: Kim, Sangwook, et al.
Veröffentlicht: (2024)
von: Kim, Sangwook, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HIME: Mitigating Object Hallucinations in LVLMs via Hallucination Insensitivity Model Editing
von: Akl, Ahmed, et al.
Veröffentlicht: (2026) -
SteerSeg: Attention Steering for Reasoning Video Segmentation
von: Cheraghian, Ali, et al.
Veröffentlicht: (2026) -
Test-Time Adaptation for Anomaly Segmentation via Topology-Aware Optimal Transport Chaining
von: Zia, Ali, et al.
Veröffentlicht: (2026) -
Questioning the Stability of Visual Question Answering
von: Rosenfeld, Amir, et al.
Veröffentlicht: (2025) -
NeuralPrefix: A Zero-shot Sensory Data Imputation Plugin
von: Khamis, Abdelwahed, et al.
Veröffentlicht: (2025)