Improving Automatic VQA Evaluation Using Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Mañas, Oscar, Krojer, Benno, Agrawal, Aishwarya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
di: Ahmadi, Saba, et al.
Pubblicazione: (2023)
di: Ahmadi, Saba, et al.
Pubblicazione: (2023)
Controlling Multimodal LLMs via Reward-guided Decoding
di: Mañas, Oscar, et al.
Pubblicazione: (2025)
di: Mañas, Oscar, et al.
Pubblicazione: (2025)
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
di: Srivastava, Archita, et al.
Pubblicazione: (2025)
di: Srivastava, Archita, et al.
Pubblicazione: (2025)
VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis
di: Mehta, Manas, et al.
Pubblicazione: (2025)
di: Mehta, Manas, et al.
Pubblicazione: (2025)
NegVQA: Can Vision Language Models Understand Negation?
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
di: Zhou, Yucheng, et al.
Pubblicazione: (2025)
di: Zhou, Yucheng, et al.
Pubblicazione: (2025)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
di: Cherian, Anoop, et al.
Pubblicazione: (2024)
di: Cherian, Anoop, et al.
Pubblicazione: (2024)
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
di: Wu, Junjie, et al.
Pubblicazione: (2024)
di: Wu, Junjie, et al.
Pubblicazione: (2024)
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
di: Srivastava, Varun, et al.
Pubblicazione: (2025)
di: Srivastava, Varun, et al.
Pubblicazione: (2025)
Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
di: Kenthapadi, Krishnaram, et al.
Pubblicazione: (2024)
di: Kenthapadi, Krishnaram, et al.
Pubblicazione: (2024)
Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras
di: Hong, Jindong, et al.
Pubblicazione: (2025)
di: Hong, Jindong, et al.
Pubblicazione: (2025)
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
Large Body Language Models
di: Punjwani, Saif, et al.
Pubblicazione: (2024)
di: Punjwani, Saif, et al.
Pubblicazione: (2024)
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
di: Krojer, Benno, et al.
Pubblicazione: (2026)
di: Krojer, Benno, et al.
Pubblicazione: (2026)
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
di: Lan, Tian, et al.
Pubblicazione: (2025)
di: Lan, Tian, et al.
Pubblicazione: (2025)
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
di: Huang, Chengyue, et al.
Pubblicazione: (2025)
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
di: Eger, Steffen, et al.
Pubblicazione: (2025)
di: Eger, Steffen, et al.
Pubblicazione: (2025)
A Survey on Multimodal Large Language Models
di: Yin, Shukang, et al.
Pubblicazione: (2023)
di: Yin, Shukang, et al.
Pubblicazione: (2023)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
di: Karim, A H M Rezaul, et al.
Pubblicazione: (2025)
di: Karim, A H M Rezaul, et al.
Pubblicazione: (2025)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
Vision-Language Models Can Self-Improve Reasoning via Reflection
di: Cheng, Kanzhi, et al.
Pubblicazione: (2024)
di: Cheng, Kanzhi, et al.
Pubblicazione: (2024)
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
di: Huang, Hailang, et al.
Pubblicazione: (2024)
di: Huang, Hailang, et al.
Pubblicazione: (2024)
Small Vision-Language Models: A Survey on Compact Architectures and Techniques
di: Patnaik, Nitesh, et al.
Pubblicazione: (2025)
di: Patnaik, Nitesh, et al.
Pubblicazione: (2025)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
di: Yu, Weihao, et al.
Pubblicazione: (2023)
di: Yu, Weihao, et al.
Pubblicazione: (2023)
Coordinated Robustness Evaluation Framework for Vision-Language Models
di: Babu, Ashwin Ramesh, et al.
Pubblicazione: (2025)
di: Babu, Ashwin Ramesh, et al.
Pubblicazione: (2025)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
di: Yin, Shukang, et al.
Pubblicazione: (2023)
di: Yin, Shukang, et al.
Pubblicazione: (2023)
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
di: Mo, Wentao, et al.
Pubblicazione: (2024)
di: Mo, Wentao, et al.
Pubblicazione: (2024)
Visual Question Decomposition on Multimodal Large Language Models
di: Zhang, Haowei, et al.
Pubblicazione: (2024)
di: Zhang, Haowei, et al.
Pubblicazione: (2024)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
di: Wan, David, et al.
Pubblicazione: (2024)
di: Wan, David, et al.
Pubblicazione: (2024)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2024)
di: Miyai, Atsuyuki, et al.
Pubblicazione: (2024)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
di: Lu, Shiyin, et al.
Pubblicazione: (2024)
di: Lu, Shiyin, et al.
Pubblicazione: (2024)
Towards Visual Text Grounding of Multimodal Large Language Model
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
di: Qin, Yulei, et al.
Pubblicazione: (2025)
di: Qin, Yulei, et al.
Pubblicazione: (2025)
Can Large Language Models Understand Symbolic Graphics Programs?
di: Qiu, Zeju, et al.
Pubblicazione: (2024)
di: Qiu, Zeju, et al.
Pubblicazione: (2024)
ARB-LLM: Alternating Refined Binarizations for Large Language Models
di: Li, Zhiteng, et al.
Pubblicazione: (2024)
di: Li, Zhiteng, et al.
Pubblicazione: (2024)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
di: Didolkar, Aniket, et al.
Pubblicazione: (2025)
di: Didolkar, Aniket, et al.
Pubblicazione: (2025)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
di: Li, Zongxia, et al.
Pubblicazione: (2025)
di: Li, Zongxia, et al.
Pubblicazione: (2025)
Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
di: Hong, Jindong, et al.
Pubblicazione: (2025)
di: Hong, Jindong, et al.
Pubblicazione: (2025)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
di: Ahmadi, Saba, et al.
Pubblicazione: (2023) -
Controlling Multimodal LLMs via Reward-guided Decoding
di: Mañas, Oscar, et al.
Pubblicazione: (2025) -
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
di: Srivastava, Archita, et al.
Pubblicazione: (2025) -
VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis
di: Mehta, Manas, et al.
Pubblicazione: (2025) -
NegVQA: Can Vision Language Models Understand Negation?
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)