GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Wenhao, Yao, Huanjin, Zhang, Mengxi, Song, Yuxin, Ouyang, Wanli, Wang, Jingdong |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dense Connector for MLLMs
by: Yao, Huanjin, et al.
Published: (2024)
by: Yao, Huanjin, et al.
Published: (2024)
GPT4Ego: Unleashing the Potential of Pre-trained Models for Zero-Shot Egocentric Action Recognition
by: Dai, Guangzhao, et al.
Published: (2024)
by: Dai, Guangzhao, et al.
Published: (2024)
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024)
by: Zhang, Mengxi, et al.
Published: (2024)
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)
by: Zhang, Jiangning, et al.
Published: (2023)
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
Zero-shot Building Age Classification from Facade Image Using GPT-4
by: Zeng, Zichao, et al.
Published: (2024)
by: Zeng, Zichao, et al.
Published: (2024)
LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?
by: Li, Bangyan, et al.
Published: (2025)
by: Li, Bangyan, et al.
Published: (2025)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
by: Zhang, Bin, et al.
Published: (2025)
by: Zhang, Bin, et al.
Published: (2025)
Can GPT-4 Models Detect Misleading Visualizations?
by: Alexander, Jason, et al.
Published: (2024)
by: Alexander, Jason, et al.
Published: (2024)
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
by: Lu, Hao, et al.
Published: (2024)
by: Lu, Hao, et al.
Published: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)
by: Xu, Siyu, et al.
Published: (2024)
MedGemma vs GPT-4: Open-Source and Proprietary Zero-shot Medical Disease Classification from Images
by: Prottasha, Md. Sazzadul Islam, et al.
Published: (2025)
by: Prottasha, Md. Sazzadul Islam, et al.
Published: (2025)
Why Compress What You Can Generate? When GPT-4o Generation Ushers in Image Compression Fields
by: Gao, Yixin, et al.
Published: (2025)
by: Gao, Yixin, et al.
Published: (2025)
An Empirical Study of GPT-4o Image Generation Capabilities
by: Chen, Sixiang, et al.
Published: (2025)
by: Chen, Sixiang, et al.
Published: (2025)
Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
by: Guo, Haonan, et al.
Published: (2024)
by: Guo, Haonan, et al.
Published: (2024)
Leveraging YOLO-World and GPT-4V LMMs for Zero-Shot Person Detection and Action Recognition in Drone Imagery
by: Limberg, Christian, et al.
Published: (2024)
by: Limberg, Christian, et al.
Published: (2024)
VGDiffZero: Text-to-image Diffusion Models Can Be Zero-shot Visual Grounders
by: Liu, Xuyang, et al.
Published: (2023)
by: Liu, Xuyang, et al.
Published: (2023)
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
by: Abdullahu, Fitim, et al.
Published: (2025)
by: Abdullahu, Fitim, et al.
Published: (2025)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
by: Wu, Yiqi, et al.
Published: (2024)
by: Wu, Yiqi, et al.
Published: (2024)
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
by: Zhang, Yaqi, et al.
Published: (2023)
by: Zhang, Yaqi, et al.
Published: (2023)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
GPTDrawer: Enhancing Visual Synthesis through ChatGPT
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
by: Xu, Yifang, et al.
Published: (2024)
by: Xu, Yifang, et al.
Published: (2024)
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
by: Zhu, Jiaqi, et al.
Published: (2024)
by: Zhu, Jiaqi, et al.
Published: (2024)
GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives
by: Chen, Zuyao, et al.
Published: (2023)
by: Chen, Zuyao, et al.
Published: (2023)
HierCode: A Lightweight Hierarchical Codebook for Zero-shot Chinese Text Recognition
by: Zhang, Yuyi, et al.
Published: (2024)
by: Zhang, Yuyi, et al.
Published: (2024)
GPT-4V Explorations: Mining Autonomous Driving
by: Li, Zixuan
Published: (2024)
by: Li, Zixuan
Published: (2024)
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning
by: Lv, Jiaxi, et al.
Published: (2023)
by: Lv, Jiaxi, et al.
Published: (2023)
GPT4Point: A Unified Framework for Point-Language Understanding and Generation
by: Qi, Zhangyang, et al.
Published: (2023)
by: Qi, Zhangyang, et al.
Published: (2023)
FG-MDM: Towards Zero-Shot Human Motion Generation via ChatGPT-Refined Descriptions
by: Shi, Xu, et al.
Published: (2023)
by: Shi, Xu, et al.
Published: (2023)
VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation
by: Deng, Jie, et al.
Published: (2026)
by: Deng, Jie, et al.
Published: (2026)
SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
by: Tiwari, Adarsh, et al.
Published: (2024)
by: Tiwari, Adarsh, et al.
Published: (2024)
TCFormer: Visual Recognition via Token Clustering Transformer
by: Zeng, Wang, et al.
Published: (2024)
by: Zeng, Wang, et al.
Published: (2024)
Exploring Visual Culture Awareness in GPT-4V: A Comprehensive Probing
by: Cao, Yong, et al.
Published: (2024)
by: Cao, Yong, et al.
Published: (2024)
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
by: Shen, Yang, et al.
Published: (2024)
by: Shen, Yang, et al.
Published: (2024)
MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4
by: Azizi, Vahid, et al.
Published: (2024)
by: Azizi, Vahid, et al.
Published: (2024)
Similar Items
-
Dense Connector for MLLMs
by: Yao, Huanjin, et al.
Published: (2024) -
GPT4Ego: Unleashing the Potential of Pre-trained Models for Zero-Shot Egocentric Action Recognition
by: Dai, Guangzhao, et al.
Published: (2024) -
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023) -
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024) -
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)