What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases
Fuente:
arXiv
Salvato in:
| Autori principali: | Tiong, Anthony Meng Huat, Zhao, Junqi, Li, Boyang, Li, Junnan, Hoi, Steven C. H., Xiong, Caiming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Are We on the Right Way for Evaluating Large Vision-Language Models?
di: Chen, Lin, et al.
Pubblicazione: (2024)
di: Chen, Lin, et al.
Pubblicazione: (2024)
What We Talk About When We Talk About Frameworks in HCI
di: Fang, Shitao, et al.
Pubblicazione: (2026)
di: Fang, Shitao, et al.
Pubblicazione: (2026)
What Do We Mean When We Talk About Data Storytelling?
di: Yang, Leni, et al.
Pubblicazione: (2025)
di: Yang, Leni, et al.
Pubblicazione: (2025)
What We Talk About When We Talk About Microbial Species
di: Apurva Narechania, et al.
Pubblicazione: (2025)
di: Apurva Narechania, et al.
Pubblicazione: (2025)
What We Talk About When We Talk About Psychometrics (in Spanish)
di: Ana R. Delgado
Pubblicazione: (2023)
di: Ana R. Delgado
Pubblicazione: (2023)
What We Talk About When We Talk About Dissipative Quantum Chaos
di: Sá, Lucas, et al.
Pubblicazione: (2026)
di: Sá, Lucas, et al.
Pubblicazione: (2026)
The Transition Issue: What Do We Change (or Not) When We Change Antimicrobial Use?
di: Nicolas Fortané, et al.
Pubblicazione: (2024)
di: Nicolas Fortané, et al.
Pubblicazione: (2024)
What We Publish and What We Do Not
di: David G. Amaral
Pubblicazione: (2025)
di: David G. Amaral
Pubblicazione: (2025)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
di: Yang, Jing, et al.
Pubblicazione: (2026)
di: Yang, Jing, et al.
Pubblicazione: (2026)
What We Talk About When We Talk About LMs: Implicit Paradigm Shifts and the Ship of Language Models
di: Zhu, Shengqi, et al.
Pubblicazione: (2024)
di: Zhu, Shengqi, et al.
Pubblicazione: (2024)
When We Get the Libraries We Want, Will We Want the Libraries We Get?
di: Seiler, Lauren, et al.
Pubblicazione: (1991)
di: Seiler, Lauren, et al.
Pubblicazione: (1991)
What Kind of Lithuania are We Fighting for When We Fight for the Lithuanian Freedom Fighters’ Memorial?
di: Rūta Statulevičiūtė-Kaučikienė
Pubblicazione: (2022)
di: Rūta Statulevičiūtė-Kaučikienė
Pubblicazione: (2022)
Networked Information: What Can We Expect and When?
di: Heterick, Robert C., Jr.
Pubblicazione: (1990)
di: Heterick, Robert C., Jr.
Pubblicazione: (1990)
Research: What We Need, and What We Get.
di: Converse, W. R.
Pubblicazione: (1984)
di: Converse, W. R.
Pubblicazione: (1984)
We Urgently Need Privilege Management in MCP: A Measurement of API Usage in MCP Ecosystems
di: Li, Zhihao, et al.
Pubblicazione: (2025)
di: Li, Zhihao, et al.
Pubblicazione: (2025)
Evaluation Criteria of Information Retrieval Systems: What We Know and What We Do Not Know
di: Hariri, Nadjla, et al.
Pubblicazione: (2014)
di: Hariri, Nadjla, et al.
Pubblicazione: (2014)
Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead
di: Aguirre, Anthony
Pubblicazione: (2023)
di: Aguirre, Anthony
Pubblicazione: (2023)
What We Augment When We Augment Visualizations: A Design Elicitation Study of How We Visually Express Data Relationships
di: Guo, Grace, et al.
Pubblicazione: (2024)
di: Guo, Grace, et al.
Pubblicazione: (2024)
When Do We Not Need Larger Vision Models?
di: Shi, Baifeng, et al.
Pubblicazione: (2024)
di: Shi, Baifeng, et al.
Pubblicazione: (2024)
Interpreting Net Survival: What We Estimate Versus What We Think We Estimate
di: Smith, Matthew J.
Pubblicazione: (2026)
di: Smith, Matthew J.
Pubblicazione: (2026)
What do We Mean by an Inclusive Pharmacology Education?
di: Jennifer Koenig, et al.
Pubblicazione: (2024)
di: Jennifer Koenig, et al.
Pubblicazione: (2024)
Can We Predict Performance of Large Models across Vision-Language Tasks?
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
di: Zhao, Qinyu, et al.
Pubblicazione: (2024)
Are LLMs Smarter Than Chimpanzees? An Evaluation on Perspective Taking and Knowledge State Estimation
di: Yang, Dingyi, et al.
Pubblicazione: (2026)
di: Yang, Dingyi, et al.
Pubblicazione: (2026)
Pulsed‐Field Ablation: What We Know and What We Don't
di: Sanghamitra Mohanty, et al.
Pubblicazione: (2025)
di: Sanghamitra Mohanty, et al.
Pubblicazione: (2025)
SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation
di: Zhang, Wenyu, et al.
Pubblicazione: (2024)
di: Zhang, Wenyu, et al.
Pubblicazione: (2024)
Reward Models Identify Consistency, Not Causality
di: Xu, Yuhui, et al.
Pubblicazione: (2025)
di: Xu, Yuhui, et al.
Pubblicazione: (2025)
What Are We About!
di: Bender, David R.
Pubblicazione: (1974)
di: Bender, David R.
Pubblicazione: (1974)
What We Need
di: Hill, Chrystle, et al.
Pubblicazione: (2008)
di: Hill, Chrystle, et al.
Pubblicazione: (2008)
What Can We Do to Help? NCLB from the Administrative Perspective
di: Baule, Steven
Pubblicazione: (2004)
di: Baule, Steven
Pubblicazione: (2004)
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
di: Peng, Jierui, et al.
Pubblicazione: (2025)
di: Peng, Jierui, et al.
Pubblicazione: (2025)
Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
di: Jeong, Daniel P., et al.
Pubblicazione: (2024)
di: Jeong, Daniel P., et al.
Pubblicazione: (2024)
Not as Simple as It Looked: Are We Concluding for Biased Arrest Practices?
di: Ozer, Murat, et al.
Pubblicazione: (2024)
di: Ozer, Murat, et al.
Pubblicazione: (2024)
“We Might Get What We Need, but Not What We Want”: Children's and Adolescents' Perceptions of Subjective Socioeconomic Status in Türkiye
di: Buse Gönül, et al.
Pubblicazione: (2026)
di: Buse Gönül, et al.
Pubblicazione: (2026)
AI Substituting Labor - What Do We Know and What Can We Know?
di: Saam, Marianne
Pubblicazione: (2023)
di: Saam, Marianne
Pubblicazione: (2023)
Chapter 8 Categorising What We Study and What We Analyse, and the Exercise of Interpretation
di: Jacobs, Dirk
Pubblicazione: (2018)
di: Jacobs, Dirk
Pubblicazione: (2018)
Women in Informal Employment: What Do We Know and What Can We Do?
di: David Kucera, et al.
Pubblicazione: (2009)
di: David Kucera, et al.
Pubblicazione: (2009)
How and Why Libraries Are Changing: What We Know and What We Need To Know.
di: Troll, Denise A.
Pubblicazione: (2002)
di: Troll, Denise A.
Pubblicazione: (2002)
Hip Microinstability—Now We Know What It Is, But What Do We Do About It?
di: Jaydeep Dhillon, et al.
Pubblicazione: (2026)
di: Jaydeep Dhillon, et al.
Pubblicazione: (2026)
What We Know and What We Don't Know About the Function of γδ T Cells
di: Immo Prinz, et al.
Pubblicazione: (2025)
di: Immo Prinz, et al.
Pubblicazione: (2025)
Correction to “The Mechanics of Cilia and Flagella: What We Know and What We Need to Know”
Pubblicazione: (2025)
Pubblicazione: (2025)
Documenti analoghi
-
Are We on the Right Way for Evaluating Large Vision-Language Models?
di: Chen, Lin, et al.
Pubblicazione: (2024) -
What We Talk About When We Talk About Frameworks in HCI
di: Fang, Shitao, et al.
Pubblicazione: (2026) -
What Do We Mean When We Talk About Data Storytelling?
di: Yang, Leni, et al.
Pubblicazione: (2025) -
What We Talk About When We Talk About Microbial Species
di: Apurva Narechania, et al.
Pubblicazione: (2025) -
What We Talk About When We Talk About Psychometrics (in Spanish)
di: Ana R. Delgado
Pubblicazione: (2023)