Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Hongjian, Ge, Yue, Ding, Qi, Liao, Yixuan, Chen, Xiaoxin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024)
by: Su, Xin, et al.
Published: (2024)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
by: Ge, Jinchao, et al.
Published: (2025)
by: Ge, Jinchao, et al.
Published: (2025)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering
by: Li, Yangfu, et al.
Published: (2025)
by: Li, Yangfu, et al.
Published: (2025)
Knowledge Generation for Zero-shot Knowledge-based VQA
by: Cao, Rui, et al.
Published: (2024)
by: Cao, Rui, et al.
Published: (2024)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
Multimodal LLM With Hierarchical Mixture-of-Experts for VQA on 3D Brain MRI
by: Vepa, Arvind Murari, et al.
Published: (2025)
by: Vepa, Arvind Murari, et al.
Published: (2025)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
by: Ma, Yunsheng, et al.
Published: (2024)
by: Ma, Yunsheng, et al.
Published: (2024)
Efficient Whole Slide Pathology VQA via Token Compression
by: Lyu, Weimin, et al.
Published: (2025)
by: Lyu, Weimin, et al.
Published: (2025)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
by: Sood, Ekta, et al.
Published: (2021)
by: Sood, Ekta, et al.
Published: (2021)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
OmniCaptioner: One Captioner to Rule Them All
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
by: Tuchinda, Pume, et al.
Published: (2025)
by: Tuchinda, Pume, et al.
Published: (2025)
IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
by: Shahgir, Haz Sameen, et al.
Published: (2024)
by: Shahgir, Haz Sameen, et al.
Published: (2024)
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
by: Zhang, Junzhe, et al.
Published: (2024)
by: Zhang, Junzhe, et al.
Published: (2024)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
by: Xu, Run, et al.
Published: (2026)
by: Xu, Run, et al.
Published: (2026)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
by: Karim, A H M Rezaul, et al.
Published: (2025)
by: Karim, A H M Rezaul, et al.
Published: (2025)
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
by: Tang, Changli, et al.
Published: (2024)
by: Tang, Changli, et al.
Published: (2024)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
by: Möller, Lucas, et al.
Published: (2024)
by: Möller, Lucas, et al.
Published: (2024)
RSCC: A Large-Scale Remote Sensing Change Caption Dataset for Disaster Events
by: Chen, Zhenyuan, et al.
Published: (2025)
by: Chen, Zhenyuan, et al.
Published: (2025)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
by: Ma, Ziyu, et al.
Published: (2024)
by: Ma, Ziyu, et al.
Published: (2024)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
by: Qi, Daiqing, et al.
Published: (2024)
by: Qi, Daiqing, et al.
Published: (2024)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
Multimodal Arabic Captioning with Interpretable Visual Concept Integration
by: Elchafei, Passant, et al.
Published: (2025)
by: Elchafei, Passant, et al.
Published: (2025)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
by: Ma, Dongsheng, et al.
Published: (2026)
by: Ma, Dongsheng, et al.
Published: (2026)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
by: Baek, Jeonghun, et al.
Published: (2025)
by: Baek, Jeonghun, et al.
Published: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
by: Chaffin, Antoine, et al.
Published: (2024)
by: Chaffin, Antoine, et al.
Published: (2024)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
by: Nguyen, Nhi Ngoc-Yen, et al.
Published: (2026)
Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches
by: Sterner, Igor, et al.
Published: (2024)
by: Sterner, Igor, et al.
Published: (2024)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
by: Jia, Yiming, et al.
Published: (2025)
by: Jia, Yiming, et al.
Published: (2025)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
by: Wu, Te-Lin, et al.
Published: (2021)
by: Wu, Te-Lin, et al.
Published: (2021)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024)
by: Guo, Jarvis, et al.
Published: (2024)
Similar Items
-
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024) -
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
by: Ge, Jinchao, et al.
Published: (2025) -
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024) -
Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering
by: Li, Yangfu, et al.
Published: (2025) -
Knowledge Generation for Zero-shot Knowledge-based VQA
by: Cao, Rui, et al.
Published: (2024)