Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yuhui, Su, Yuchang, Liu, Yiming, Wang, Xiaohan, Burgess, James, Sui, Elaine, Wang, Chenyu, Aklilu, Josiah, Lozano, Alejandro, Wei, Anjiang, Schmidt, Ludwig, Yeung-Levy, Serena |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
di: Aklilu, Josiah, et al.
Pubblicazione: (2024)
di: Aklilu, Josiah, et al.
Pubblicazione: (2024)
NegVQA: Can Vision Language Models Understand Negation?
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
di: Sui, Elaine, et al.
Pubblicazione: (2024)
di: Sui, Elaine, et al.
Pubblicazione: (2024)
Why are Visually-Grounded Language Models Bad at Image Classification?
di: Zhang, Yuhui, et al.
Pubblicazione: (2024)
di: Zhang, Yuhui, et al.
Pubblicazione: (2024)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
di: Zhang, Yuhui, et al.
Pubblicazione: (2024)
di: Zhang, Yuhui, et al.
Pubblicazione: (2024)
Revisiting Active Learning in the Era of Vision Foundation Models
di: Gupte, Sanket Rajan, et al.
Pubblicazione: (2024)
di: Gupte, Sanket Rajan, et al.
Pubblicazione: (2024)
Data or Language Supervision: What Makes CLIP Better than DINO?
di: Liu, Yiming, et al.
Pubblicazione: (2025)
di: Liu, Yiming, et al.
Pubblicazione: (2025)
Closing the Modality Gap for Mixed Modality Search
di: Li, Binxu, et al.
Pubblicazione: (2025)
di: Li, Binxu, et al.
Pubblicazione: (2025)
Depth-guided NeRF Training via Earth Mover's Distance
di: Rau, Anita, et al.
Pubblicazione: (2024)
di: Rau, Anita, et al.
Pubblicazione: (2024)
V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
di: Tang, Bingda, et al.
Pubblicazione: (2026)
di: Tang, Bingda, et al.
Pubblicazione: (2026)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
di: Wang, Xiaohan, et al.
Pubblicazione: (2024)
di: Wang, Xiaohan, et al.
Pubblicazione: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
di: Endo, Mark, et al.
Pubblicazione: (2024)
di: Endo, Mark, et al.
Pubblicazione: (2024)
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
di: Lozano, Alejandro, et al.
Pubblicazione: (2025)
di: Lozano, Alejandro, et al.
Pubblicazione: (2025)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
di: Li, He, et al.
Pubblicazione: (2026)
di: Li, He, et al.
Pubblicazione: (2026)
Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
di: Wu, Shengguang, et al.
Pubblicazione: (2025)
di: Wu, Shengguang, et al.
Pubblicazione: (2025)
Video Action Differencing
di: Burgess, James, et al.
Pubblicazione: (2025)
di: Burgess, James, et al.
Pubblicazione: (2025)
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
di: Lozano, Alejandro, et al.
Pubblicazione: (2024)
di: Lozano, Alejandro, et al.
Pubblicazione: (2024)
CellFlux: Simulating Cellular Morphology Changes via Flow Matching
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
di: Zhang, Yuhui, et al.
Pubblicazione: (2025)
Temporal Preference Optimization for Long-Form Video Understanding
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
Can Large Language Models Match the Conclusions of Systematic Reviews?
di: Polzak, Christopher, et al.
Pubblicazione: (2025)
di: Polzak, Christopher, et al.
Pubblicazione: (2025)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
di: Burgess, James, et al.
Pubblicazione: (2023)
di: Burgess, James, et al.
Pubblicazione: (2023)
Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
di: Rau, Anita, et al.
Pubblicazione: (2025)
di: Rau, Anita, et al.
Pubblicazione: (2025)
A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
di: Lozano, Alejandro, et al.
Pubblicazione: (2025)
di: Lozano, Alejandro, et al.
Pubblicazione: (2025)
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
di: Sun, Min Woo, et al.
Pubblicazione: (2025)
di: Sun, Min Woo, et al.
Pubblicazione: (2025)
RadDiff: Describing Differences in Radiology Image Sets with Natural Language
di: Shen, Xiaoxian, et al.
Pubblicazione: (2026)
di: Shen, Xiaoxian, et al.
Pubblicazione: (2026)
Describing Differences in Image Sets with Natural Language
di: Dunlap, Lisa, et al.
Pubblicazione: (2023)
di: Dunlap, Lisa, et al.
Pubblicazione: (2023)
Tool Verification for Test-Time Reinforcement Learning
di: Liao, Ruotong, et al.
Pubblicazione: (2026)
di: Liao, Ruotong, et al.
Pubblicazione: (2026)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
di: Heo, Jaewoo, et al.
Pubblicazione: (2024)
di: Heo, Jaewoo, et al.
Pubblicazione: (2024)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
di: Zohar, Orr, et al.
Pubblicazione: (2024)
di: Zohar, Orr, et al.
Pubblicazione: (2024)
Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
di: Weng, Zhenzhen, et al.
Pubblicazione: (2023)
di: Weng, Zhenzhen, et al.
Pubblicazione: (2023)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
di: Burgess, James, et al.
Pubblicazione: (2026)
di: Burgess, James, et al.
Pubblicazione: (2026)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
di: Endo, Mark, et al.
Pubblicazione: (2025)
di: Endo, Mark, et al.
Pubblicazione: (2025)
Differentiating Choices via Commonality for Multiple-Choice Question Answering
di: Deng, Wenqing, et al.
Pubblicazione: (2024)
di: Deng, Wenqing, et al.
Pubblicazione: (2024)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
di: Wang, Zeyu, et al.
Pubblicazione: (2024)
di: Wang, Zeyu, et al.
Pubblicazione: (2024)
Multi-Human Mesh Recovery with Transformers
di: Wang, Zeyu, et al.
Pubblicazione: (2024)
di: Wang, Zeyu, et al.
Pubblicazione: (2024)
Qworld: Question-Specific Evaluation Criteria for LLMs
di: Gao, Shanghua, et al.
Pubblicazione: (2026)
di: Gao, Shanghua, et al.
Pubblicazione: (2026)
Multiple-Choice Questions are Efficient and Robust LLM Evaluators
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
di: Gu, Jeffrey, et al.
Pubblicazione: (2025)
di: Gu, Jeffrey, et al.
Pubblicazione: (2025)
Automated Generation of Multiple-Choice Cloze Questions for Assessing English Vocabulary Using GPT-turbo 3.5
di: Wang, Qiao, et al.
Pubblicazione: (2024)
di: Wang, Qiao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
di: Aklilu, Josiah, et al.
Pubblicazione: (2024) -
NegVQA: Can Vision Language Models Understand Negation?
di: Zhang, Yuhui, et al.
Pubblicazione: (2025) -
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
di: Sui, Elaine, et al.
Pubblicazione: (2024) -
Why are Visually-Grounded Language Models Bad at Image Classification?
di: Zhang, Yuhui, et al.
Pubblicazione: (2024) -
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
di: Zhang, Yuhui, et al.
Pubblicazione: (2024)