Gespeichert in:
| Hauptverfasser: | Pramanick, Shraman, Chellappa, Rama, Venugopalan, Subhashini |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2407.09413 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
von: Pramanick, Shraman, et al.
Veröffentlicht: (2023)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2023)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision Models
von: Wei, Guoyizhe, et al.
Veröffentlicht: (2025)
von: Wei, Guoyizhe, et al.
Veröffentlicht: (2025)
DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning
von: Yilmaz, Abdurrahim, et al.
Veröffentlicht: (2026)
von: Yilmaz, Abdurrahim, et al.
Veröffentlicht: (2026)
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding
von: Li, Zekun, et al.
Veröffentlicht: (2024)
von: Li, Zekun, et al.
Veröffentlicht: (2024)
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
von: Pang, Wei, et al.
Veröffentlicht: (2025)
von: Pang, Wei, et al.
Veröffentlicht: (2025)
ChartInsights: Evaluating Multimodal Large Language Models for Low-Level Chart Question Answering
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
Synthetic Document Question Answering in Hungarian
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
von: Butsanets, Léo, et al.
Veröffentlicht: (2025)
von: Butsanets, Léo, et al.
Veröffentlicht: (2025)
MMToM-QA: Multimodal Theory of Mind Question Answering
von: Jin, Chuanyang, et al.
Veröffentlicht: (2024)
von: Jin, Chuanyang, et al.
Veröffentlicht: (2024)
Hallucination Benchmark in Medical Visual Question Answering
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
von: Cocchi, Federico, et al.
Veröffentlicht: (2024)
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
von: Kim, Wonjoong, et al.
Veröffentlicht: (2024)
von: Kim, Wonjoong, et al.
Veröffentlicht: (2024)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
von: Kendre, Shrikant, et al.
Veröffentlicht: (2025)
von: Kendre, Shrikant, et al.
Veröffentlicht: (2025)
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
von: Srivastava, Varun, et al.
Veröffentlicht: (2025)
von: Srivastava, Varun, et al.
Veröffentlicht: (2025)
Top-down Activity Representation Learning for Video Question Answering
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
von: Yu, Wenhan, et al.
Veröffentlicht: (2025)
von: Yu, Wenhan, et al.
Veröffentlicht: (2025)
Multi-object event graph representation learning for Video Question Answering
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
von: Fu, Xingyu, et al.
Veröffentlicht: (2023)
von: Fu, Xingyu, et al.
Veröffentlicht: (2023)
DiffRegCD: Integrated Registration and Change Detection with Diffusion Features
von: Madani, Seyedehanita, et al.
Veröffentlicht: (2025)
von: Madani, Seyedehanita, et al.
Veröffentlicht: (2025)
Scientific Reasoning: Assessment of Multimodal Generative LLMs
von: Dreyer, Florian, et al.
Veröffentlicht: (2025)
von: Dreyer, Florian, et al.
Veröffentlicht: (2025)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
MediFact at MEDIQA-M3G 2024: Medical Question Answering in Dermatology with Multimodal Learning
von: Saeed, Nadia
Veröffentlicht: (2024)
von: Saeed, Nadia
Veröffentlicht: (2024)
Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations
von: Yeh, Yahsin, et al.
Veröffentlicht: (2025)
von: Yeh, Yahsin, et al.
Veröffentlicht: (2025)
UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation
von: Ghosh, Shiv, et al.
Veröffentlicht: (2026)
von: Ghosh, Shiv, et al.
Veröffentlicht: (2026)
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
von: Winata, Genta Indra, et al.
Veröffentlicht: (2024)
von: Winata, Genta Indra, et al.
Veröffentlicht: (2024)
Exploring Diverse Methods in Visual Question Answering
von: Li, Panfeng, et al.
Veröffentlicht: (2024)
von: Li, Panfeng, et al.
Veröffentlicht: (2024)
SciMDR: Advancing Scientific Multimodal Document Reasoning
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
Federated Document Visual Question Answering: A Pilot Study
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
von: Liu, Dongqi, et al.
Veröffentlicht: (2025)
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
von: Compagnoni, Alberto, et al.
Veröffentlicht: (2025)
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
von: Han, Wei, et al.
Veröffentlicht: (2023)
von: Han, Wei, et al.
Veröffentlicht: (2023)
CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering
von: Kaur, Rachneet, et al.
Veröffentlicht: (2025)
von: Kaur, Rachneet, et al.
Veröffentlicht: (2025)
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
von: Pramanick, Shraman, et al.
Veröffentlicht: (2023) -
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025) -
ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision Models
von: Wei, Guoyizhe, et al.
Veröffentlicht: (2025) -
DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning
von: Yilmaz, Abdurrahim, et al.
Veröffentlicht: (2026) -
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
von: Wang, Haibo, et al.
Veröffentlicht: (2024)