BloomVQA: Assessing Hierarchical Multi-modal Comprehension
Fuente:
arXiv
Guardado en:
| Autores principales: | Gong, Yunye, Shrestha, Robik, Claypoole, Jared, Cogswell, Michael, Ray, Arijit, Kanan, Christopher, Divakaran, Ajay |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Document Understanding, Measurement, and Manipulation Using Category Theory
por: Claypoole, Jared, et al.
Publicado: (2025)
por: Claypoole, Jared, et al.
Publicado: (2025)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
por: Shrestha, Robik, et al.
Publicado: (2020)
por: Shrestha, Robik, et al.
Publicado: (2020)
Punching Bag vs. Punching Person: Motion Transferability in Videos
por: Abdullah, Raiyaan, et al.
Publicado: (2025)
por: Abdullah, Raiyaan, et al.
Publicado: (2025)
Improving Multimodal Large Language Models Using Continual Learning
por: Srivastava, Shikhar, et al.
Publicado: (2024)
por: Srivastava, Shikhar, et al.
Publicado: (2024)
Probing Conceptual Understanding of Large Visual-Language Models
por: Schiappa, Madeline, et al.
Publicado: (2023)
por: Schiappa, Madeline, et al.
Publicado: (2023)
Revisiting Multi-Modal LLM Evaluation
por: Lu, Jian, et al.
Publicado: (2024)
por: Lu, Jian, et al.
Publicado: (2024)
Are Bias Mitigation Techniques for Deep Learning Effective?
por: Shrestha, Robik, et al.
Publicado: (2021)
por: Shrestha, Robik, et al.
Publicado: (2021)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
por: Chen, Yangyi, et al.
Publicado: (2023)
por: Chen, Yangyi, et al.
Publicado: (2023)
DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback
por: Chen, Yangyi, et al.
Publicado: (2023)
por: Chen, Yangyi, et al.
Publicado: (2023)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
por: Gwilliam, Matthew, et al.
Publicado: (2023)
por: Gwilliam, Matthew, et al.
Publicado: (2023)
Multimodal LLM With Hierarchical Mixture-of-Experts for VQA on 3D Brain MRI
por: Vepa, Arvind Murari, et al.
Publicado: (2025)
por: Vepa, Arvind Murari, et al.
Publicado: (2025)
Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing
por: Wijegunarathna, Kalana, et al.
Publicado: (2025)
por: Wijegunarathna, Kalana, et al.
Publicado: (2025)
GRAM: Global Reasoning for Multi-Page VQA
por: Blau, Tsachi, et al.
Publicado: (2024)
por: Blau, Tsachi, et al.
Publicado: (2024)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
por: Guo, Ziyu, et al.
Publicado: (2025)
por: Guo, Ziyu, et al.
Publicado: (2025)
CommVQA: Situating Visual Question Answering in Communicative Contexts
por: Naik, Nandita Shankar, et al.
Publicado: (2024)
por: Naik, Nandita Shankar, et al.
Publicado: (2024)
Pelican: Correcting Hallucination in Vision-LLMs via Claim Decomposition and Program of Thought Verification
por: Sahu, Pritish, et al.
Publicado: (2024)
por: Sahu, Pritish, et al.
Publicado: (2024)
DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
por: Liu, Zixuan, et al.
Publicado: (2025)
por: Liu, Zixuan, et al.
Publicado: (2025)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
por: Fu, Chaoyou, et al.
Publicado: (2024)
por: Fu, Chaoyou, et al.
Publicado: (2024)
GRASP: A Rehearsal Policy for Efficient Online Continual Learning
por: Harun, Md Yousuf, et al.
Publicado: (2023)
por: Harun, Md Yousuf, et al.
Publicado: (2023)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
por: Yu, Suhao, et al.
Publicado: (2025)
por: Yu, Suhao, et al.
Publicado: (2025)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
por: Zhang, Kaichen, et al.
Publicado: (2024)
por: Zhang, Kaichen, et al.
Publicado: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
por: Zhang, Ming, et al.
Publicado: (2024)
por: Zhang, Ming, et al.
Publicado: (2024)
Attribute Diversity Determines the Systematicity Gap in VQA
por: Berlot-Attwell, Ian, et al.
Publicado: (2023)
por: Berlot-Attwell, Ian, et al.
Publicado: (2023)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
por: Fu, Jinlan, et al.
Publicado: (2025)
por: Fu, Jinlan, et al.
Publicado: (2025)
Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
por: Madaan, Divyam, et al.
Publicado: (2025)
por: Madaan, Divyam, et al.
Publicado: (2025)
OccamNets: Mitigating Dataset Bias by Favoring Simpler Hypotheses
por: Shrestha, Robik, et al.
Publicado: (2022)
por: Shrestha, Robik, et al.
Publicado: (2022)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
por: Qin, Lixiong, et al.
Publicado: (2025)
por: Qin, Lixiong, et al.
Publicado: (2025)
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
por: Wu, Chengyue, et al.
Publicado: (2024)
por: Wu, Chengyue, et al.
Publicado: (2024)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
por: Rostamkhani, Mohammadmostafa, et al.
Publicado: (2024)
por: Rostamkhani, Mohammadmostafa, et al.
Publicado: (2024)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
por: Ma, Dongsheng, et al.
Publicado: (2026)
por: Ma, Dongsheng, et al.
Publicado: (2026)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
por: Liang, Jianxin, et al.
Publicado: (2025)
por: Liang, Jianxin, et al.
Publicado: (2025)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
por: Ge, Jinchao, et al.
Publicado: (2025)
por: Ge, Jinchao, et al.
Publicado: (2025)
Efficient Whole Slide Pathology VQA via Token Compression
por: Lyu, Weimin, et al.
Publicado: (2025)
por: Lyu, Weimin, et al.
Publicado: (2025)
MINDS: A Cross-cultural Dialogue Corpus for Social Norm Classification and Adherence Detection
por: Sahu, Pritish, et al.
Publicado: (2025)
por: Sahu, Pritish, et al.
Publicado: (2025)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
por: Yang, Rui, et al.
Publicado: (2025)
por: Yang, Rui, et al.
Publicado: (2025)
Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability
por: Yoon, Yejun, et al.
Publicado: (2024)
por: Yoon, Yejun, et al.
Publicado: (2024)
Instruct-Imagen: Image Generation with Multi-modal Instruction
por: Hu, Hexiang, et al.
Publicado: (2024)
por: Hu, Hexiang, et al.
Publicado: (2024)
Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches
por: Sterner, Igor, et al.
Publicado: (2024)
por: Sterner, Igor, et al.
Publicado: (2024)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
por: Hou, Wenjun, et al.
Publicado: (2024)
por: Hou, Wenjun, et al.
Publicado: (2024)
IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
por: Shahgir, Haz Sameen, et al.
Publicado: (2024)
por: Shahgir, Haz Sameen, et al.
Publicado: (2024)
Ejemplares similares
-
Document Understanding, Measurement, and Manipulation Using Category Theory
por: Claypoole, Jared, et al.
Publicado: (2025) -
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
por: Shrestha, Robik, et al.
Publicado: (2020) -
Punching Bag vs. Punching Person: Motion Transferability in Videos
por: Abdullah, Raiyaan, et al.
Publicado: (2025) -
Improving Multimodal Large Language Models Using Continual Learning
por: Srivastava, Shikhar, et al.
Publicado: (2024) -
Probing Conceptual Understanding of Large Visual-Language Models
por: Schiappa, Madeline, et al.
Publicado: (2023)