How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Ramachandran, Rahul, Garjani, Ali, Bachmann, Roman, Atanov, Andrei, Kar, Oğuzhan Fatih, Zamir, Amir |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
di: Bachmann, Roman, et al.
Pubblicazione: (2024)
di: Bachmann, Roman, et al.
Pubblicazione: (2024)
Unraveling the Key Components of OOD Generalization via Diversification
di: Benoit, Harold, et al.
Pubblicazione: (2023)
di: Benoit, Harold, et al.
Pubblicazione: (2023)
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
di: Atanov, Andrei, et al.
Pubblicazione: (2026)
di: Atanov, Andrei, et al.
Pubblicazione: (2026)
Solving Vision Tasks with Simple Photoreceptors Instead of Cameras
di: Atanov, Andrei, et al.
Pubblicazione: (2024)
di: Atanov, Andrei, et al.
Pubblicazione: (2024)
Large (Vision) Language Models are Unsupervised In-Context Learners
di: Gadetsky, Artyom, et al.
Pubblicazione: (2025)
di: Gadetsky, Artyom, et al.
Pubblicazione: (2025)
(1D) Ordered Tokens Enable Efficient Test-Time Search
di: Gao, Zhitong, et al.
Pubblicazione: (2026)
di: Gao, Zhitong, et al.
Pubblicazione: (2026)
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
di: Bachmann, Roman, et al.
Pubblicazione: (2025)
di: Bachmann, Roman, et al.
Pubblicazione: (2025)
BRAVE: Broadening the visual encoding of vision-language models
di: Kar, Oğuzhan Fatih, et al.
Pubblicazione: (2024)
di: Kar, Oğuzhan Fatih, et al.
Pubblicazione: (2024)
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
di: Kar, Oğuzhan Fatih, et al.
Pubblicazione: (2026)
di: Kar, Oğuzhan Fatih, et al.
Pubblicazione: (2026)
Controlled Training Data Generation with Diffusion Models
di: Yeo, Teresa, et al.
Pubblicazione: (2024)
di: Yeo, Teresa, et al.
Pubblicazione: (2024)
ViPer: Visual Personalization of Generative Models via Individual Preference Learning
di: Salehi, Sogand, et al.
Pubblicazione: (2024)
di: Salehi, Sogand, et al.
Pubblicazione: (2024)
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
di: Kelly, Chris, et al.
Pubblicazione: (2024)
di: Kelly, Chris, et al.
Pubblicazione: (2024)
Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
di: Bourceanu, Radu-Andrei, et al.
Pubblicazione: (2025)
di: Bourceanu, Radu-Andrei, et al.
Pubblicazione: (2025)
On Evaluation of Vision Datasets and Models using Human Competency Frameworks
di: Ramachandran, Rahul, et al.
Pubblicazione: (2024)
di: Ramachandran, Rahul, et al.
Pubblicazione: (2024)
Understanding AI Evaluation Patterns: How Different GPT Models Assess Vision-Language Descriptions
di: Abdoli, Sajjad, et al.
Pubblicazione: (2025)
di: Abdoli, Sajjad, et al.
Pubblicazione: (2025)
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
di: Kelly, Chris, et al.
Pubblicazione: (2024)
di: Kelly, Chris, et al.
Pubblicazione: (2024)
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
di: Zhang, Kai, et al.
Pubblicazione: (2023)
di: Zhang, Kai, et al.
Pubblicazione: (2023)
Editorial Comment on (Abiraterone Acetate Triggers ER Stress‐Mediated Androgen Receptor Suppression via PERK / ATF4 / CHOP Signaling in Prostate Cancer)
di: Fatih Kar
Pubblicazione: (2026)
di: Fatih Kar
Pubblicazione: (2026)
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
di: Truong, Thanh-Dat, et al.
Pubblicazione: (2025)
di: Truong, Thanh-Dat, et al.
Pubblicazione: (2025)
FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing
di: Corley, Isaac, et al.
Pubblicazione: (2025)
di: Corley, Isaac, et al.
Pubblicazione: (2025)
Vision Foundation Models for Computed Tomography
di: Pai, Suraj, et al.
Pubblicazione: (2025)
di: Pai, Suraj, et al.
Pubblicazione: (2025)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
A Systematic Literature Review on Deep Learning-based Depth Estimation in Computer Vision
di: Rohan, Ali, et al.
Pubblicazione: (2025)
di: Rohan, Ali, et al.
Pubblicazione: (2025)
Understanding the Transfer Limits of Vision Foundation Models
di: Huang, Shiqi, et al.
Pubblicazione: (2026)
di: Huang, Shiqi, et al.
Pubblicazione: (2026)
Evaluating ChatGPT-4 Vision on Brazil's National Undergraduate Computer Science Exam
di: Mendonça, Nabor C.
Pubblicazione: (2024)
di: Mendonça, Nabor C.
Pubblicazione: (2024)
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
di: Brusnicki, Roberto, et al.
Pubblicazione: (2026)
di: Brusnicki, Roberto, et al.
Pubblicazione: (2026)
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
di: Yang, Bo, et al.
Pubblicazione: (2025)
di: Yang, Bo, et al.
Pubblicazione: (2025)
RegionGPT: Towards Region Understanding Vision Language Model
di: Guo, Qiushan, et al.
Pubblicazione: (2024)
di: Guo, Qiushan, et al.
Pubblicazione: (2024)
How Does Diverse Interpretability of Textual Prompts Impact Medical Vision-Language Zero-Shot Tasks?
di: Wang, Sicheng, et al.
Pubblicazione: (2024)
di: Wang, Sicheng, et al.
Pubblicazione: (2024)
Understanding Task Transfer in Vision-Language Models
di: Sachdeva, Bhuvan, et al.
Pubblicazione: (2025)
di: Sachdeva, Bhuvan, et al.
Pubblicazione: (2025)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
di: Ashraf, Tajamul, et al.
Pubblicazione: (2025)
di: Ashraf, Tajamul, et al.
Pubblicazione: (2025)
Task-Agnostic Attacks Against Vision Foundation Models
di: Pulfer, Brian, et al.
Pubblicazione: (2025)
di: Pulfer, Brian, et al.
Pubblicazione: (2025)
GazeVLM: A Vision-Language Model for Multi-Task Gaze Understanding
di: Mathew, Athul M., et al.
Pubblicazione: (2025)
di: Mathew, Athul M., et al.
Pubblicazione: (2025)
How Well Does GPT-4V(ision) Adapt to Distribution Shifts? A Preliminary Investigation
di: Han, Zhongyi, et al.
Pubblicazione: (2023)
di: Han, Zhongyi, et al.
Pubblicazione: (2023)
VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning
di: Song, Yunpeng, et al.
Pubblicazione: (2023)
di: Song, Yunpeng, et al.
Pubblicazione: (2023)
Tensor Train Decomposition for Adversarial Attacks on Computer Vision Models
di: Chertkov, Andrei, et al.
Pubblicazione: (2023)
di: Chertkov, Andrei, et al.
Pubblicazione: (2023)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
di: Wang, Hanbin, et al.
Pubblicazione: (2025)
di: Wang, Hanbin, et al.
Pubblicazione: (2025)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
di: Meng, Fanqing, et al.
Pubblicazione: (2024)
di: Meng, Fanqing, et al.
Pubblicazione: (2024)
A Multimodal Vision Foundation Model for Clinical Dermatology
di: Yan, Siyuan, et al.
Pubblicazione: (2024)
di: Yan, Siyuan, et al.
Pubblicazione: (2024)
Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency
di: Shahriar, Sakib, et al.
Pubblicazione: (2024)
di: Shahriar, Sakib, et al.
Pubblicazione: (2024)
Documenti analoghi
-
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
di: Bachmann, Roman, et al.
Pubblicazione: (2024) -
Unraveling the Key Components of OOD Generalization via Diversification
di: Benoit, Harold, et al.
Pubblicazione: (2023) -
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
di: Atanov, Andrei, et al.
Pubblicazione: (2026) -
Solving Vision Tasks with Simple Photoreceptors Instead of Cameras
di: Atanov, Andrei, et al.
Pubblicazione: (2024) -
Large (Vision) Language Models are Unsupervised In-Context Learners
di: Gadetsky, Artyom, et al.
Pubblicazione: (2025)