Salvato in:
| Autori principali: | Wu, Wenhao, Yao, Huanjin, Zhang, Mengxi, Song, Yuxin, Ouyang, Wanli, Wang, Jingdong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2311.15732 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dense Connector for MLLMs
di: Yao, Huanjin, et al.
Pubblicazione: (2024)
di: Yao, Huanjin, et al.
Pubblicazione: (2024)
Automated Multi-level Preference for MLLMs
di: Zhang, Mengxi, et al.
Pubblicazione: (2024)
di: Zhang, Mengxi, et al.
Pubblicazione: (2024)
GPT4Ego: Unleashing the Potential of Pre-trained Models for Zero-Shot Egocentric Action Recognition
di: Dai, Guangzhao, et al.
Pubblicazione: (2024)
di: Dai, Guangzhao, et al.
Pubblicazione: (2024)
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
di: Lian, Zheng, et al.
Pubblicazione: (2023)
di: Lian, Zheng, et al.
Pubblicazione: (2023)
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
di: Zhang, Jiangning, et al.
Pubblicazione: (2023)
di: Zhang, Jiangning, et al.
Pubblicazione: (2023)
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding
di: Sun, Qi, et al.
Pubblicazione: (2024)
di: Sun, Qi, et al.
Pubblicazione: (2024)
LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?
di: Li, Bangyan, et al.
Pubblicazione: (2025)
di: Li, Bangyan, et al.
Pubblicazione: (2025)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
di: Zhang, Sha, et al.
Pubblicazione: (2024)
di: Zhang, Sha, et al.
Pubblicazione: (2024)
Zero-shot Building Age Classification from Facade Image Using GPT-4
di: Zeng, Zichao, et al.
Pubblicazione: (2024)
di: Zeng, Zichao, et al.
Pubblicazione: (2024)
Can GPT-4 Models Detect Misleading Visualizations?
di: Alexander, Jason, et al.
Pubblicazione: (2024)
di: Alexander, Jason, et al.
Pubblicazione: (2024)
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
di: Zhang, Bin, et al.
Pubblicazione: (2025)
di: Zhang, Bin, et al.
Pubblicazione: (2025)
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
di: Lu, Hao, et al.
Pubblicazione: (2024)
di: Lu, Hao, et al.
Pubblicazione: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
di: Xu, Siyu, et al.
Pubblicazione: (2024)
di: Xu, Siyu, et al.
Pubblicazione: (2024)
MedGemma vs GPT-4: Open-Source and Proprietary Zero-shot Medical Disease Classification from Images
di: Prottasha, Md. Sazzadul Islam, et al.
Pubblicazione: (2025)
di: Prottasha, Md. Sazzadul Islam, et al.
Pubblicazione: (2025)
An Empirical Study of GPT-4o Image Generation Capabilities
di: Chen, Sixiang, et al.
Pubblicazione: (2025)
di: Chen, Sixiang, et al.
Pubblicazione: (2025)
Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
di: Guo, Haonan, et al.
Pubblicazione: (2024)
di: Guo, Haonan, et al.
Pubblicazione: (2024)
Why Compress What You Can Generate? When GPT-4o Generation Ushers in Image Compression Fields
di: Gao, Yixin, et al.
Pubblicazione: (2025)
di: Gao, Yixin, et al.
Pubblicazione: (2025)
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
di: Zhang, Yaqi, et al.
Pubblicazione: (2023)
di: Zhang, Yaqi, et al.
Pubblicazione: (2023)
Leveraging YOLO-World and GPT-4V LMMs for Zero-Shot Person Detection and Action Recognition in Drone Imagery
di: Limberg, Christian, et al.
Pubblicazione: (2024)
di: Limberg, Christian, et al.
Pubblicazione: (2024)
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
di: Singh, Simranjit, et al.
Pubblicazione: (2024)
di: Singh, Simranjit, et al.
Pubblicazione: (2024)
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
di: Wang, Yuan, et al.
Pubblicazione: (2024)
di: Wang, Yuan, et al.
Pubblicazione: (2024)
VGDiffZero: Text-to-image Diffusion Models Can Be Zero-shot Visual Grounders
di: Liu, Xuyang, et al.
Pubblicazione: (2023)
di: Liu, Xuyang, et al.
Pubblicazione: (2023)
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
di: Yao, Huanjin, et al.
Pubblicazione: (2024)
di: Yao, Huanjin, et al.
Pubblicazione: (2024)
CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation
di: Song, YuXin, et al.
Pubblicazione: (2026)
di: Song, YuXin, et al.
Pubblicazione: (2026)
GPTDrawer: Enhancing Visual Synthesis through ChatGPT
di: Li, Kun, et al.
Pubblicazione: (2024)
di: Li, Kun, et al.
Pubblicazione: (2024)
Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
di: Abdullahu, Fitim, et al.
Pubblicazione: (2025)
di: Abdullahu, Fitim, et al.
Pubblicazione: (2025)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
di: Wu, Yiqi, et al.
Pubblicazione: (2024)
di: Wu, Yiqi, et al.
Pubblicazione: (2024)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT
di: Xu, Yifang, et al.
Pubblicazione: (2024)
di: Xu, Yifang, et al.
Pubblicazione: (2024)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
di: Yan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Yan, Zhiyuan, et al.
Pubblicazione: (2025)
TCFormer: Visual Recognition via Token Clustering Transformer
di: Zeng, Wang, et al.
Pubblicazione: (2024)
di: Zeng, Wang, et al.
Pubblicazione: (2024)
Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection
di: Zhu, Jiaqi, et al.
Pubblicazione: (2024)
di: Zhu, Jiaqi, et al.
Pubblicazione: (2024)
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
di: Shen, Yang, et al.
Pubblicazione: (2024)
di: Shen, Yang, et al.
Pubblicazione: (2024)
HierCode: A Lightweight Hierarchical Codebook for Zero-shot Chinese Text Recognition
di: Zhang, Yuyi, et al.
Pubblicazione: (2024)
di: Zhang, Yuyi, et al.
Pubblicazione: (2024)
VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation
di: Deng, Jie, et al.
Pubblicazione: (2026)
di: Deng, Jie, et al.
Pubblicazione: (2026)
FG-MDM: Towards Zero-Shot Human Motion Generation via ChatGPT-Refined Descriptions
di: Shi, Xu, et al.
Pubblicazione: (2023)
di: Shi, Xu, et al.
Pubblicazione: (2023)
GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives
di: Chen, Zuyao, et al.
Pubblicazione: (2023)
di: Chen, Zuyao, et al.
Pubblicazione: (2023)
Anomagic: Crossmodal Prompt-driven Zero-shot Anomaly Generation
di: Jiang, Yuxin, et al.
Pubblicazione: (2025)
di: Jiang, Yuxin, et al.
Pubblicazione: (2025)
SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
di: Tiwari, Adarsh, et al.
Pubblicazione: (2024)
di: Tiwari, Adarsh, et al.
Pubblicazione: (2024)
MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4
di: Azizi, Vahid, et al.
Pubblicazione: (2024)
di: Azizi, Vahid, et al.
Pubblicazione: (2024)
SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
di: Jin, Zhao, et al.
Pubblicazione: (2025)
di: Jin, Zhao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Dense Connector for MLLMs
di: Yao, Huanjin, et al.
Pubblicazione: (2024) -
Automated Multi-level Preference for MLLMs
di: Zhang, Mengxi, et al.
Pubblicazione: (2024) -
GPT4Ego: Unleashing the Potential of Pre-trained Models for Zero-Shot Egocentric Action Recognition
di: Dai, Guangzhao, et al.
Pubblicazione: (2024) -
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
di: Lian, Zheng, et al.
Pubblicazione: (2023) -
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
di: Zhang, Jiangning, et al.
Pubblicazione: (2023)