GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Kaichun, Chen, Jian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GesGPT: Speech Gesture Synthesis With Text Parsing from ChatGPT
di: Gao, Nan, et al.
Pubblicazione: (2023)
di: Gao, Nan, et al.
Pubblicazione: (2023)
ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
di: Xia, Ding, et al.
Pubblicazione: (2025)
di: Xia, Ding, et al.
Pubblicazione: (2025)
Human-Centred Evaluation of Text-to-Image Generation Models for Self-expression of Mental Distress: A Dataset Based on GPT-4o
di: He, Sui, et al.
Pubblicazione: (2025)
di: He, Sui, et al.
Pubblicazione: (2025)
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
Can ChatGPT assist visually impaired people with micro-navigation?
di: He, Junxian, et al.
Pubblicazione: (2024)
di: He, Junxian, et al.
Pubblicazione: (2024)
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos
di: Huang, Xiyang, et al.
Pubblicazione: (2026)
di: Huang, Xiyang, et al.
Pubblicazione: (2026)
GazeGPT: Augmenting Human Capabilities using Gaze-contingent Contextual AI for Smart Eyewear
di: Konrad, Robert, et al.
Pubblicazione: (2024)
di: Konrad, Robert, et al.
Pubblicazione: (2024)
AltChart: Enhancing VLM-based Chart Summarization Through Multi-Pretext Tasks
di: Moured, Omar, et al.
Pubblicazione: (2024)
di: Moured, Omar, et al.
Pubblicazione: (2024)
A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
di: Betala, Siddharth, et al.
Pubblicazione: (2025)
di: Betala, Siddharth, et al.
Pubblicazione: (2025)
Alt4Blind: A User Interface to Simplify Charts Alt-Text Creation
di: Moured, Omar, et al.
Pubblicazione: (2024)
di: Moured, Omar, et al.
Pubblicazione: (2024)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
di: Wu, Zhiyong, et al.
Pubblicazione: (2024)
di: Wu, Zhiyong, et al.
Pubblicazione: (2024)
A Review on Large Language Models for Visual Analytics
di: Agarwal, Navya Sonal, et al.
Pubblicazione: (2025)
di: Agarwal, Navya Sonal, et al.
Pubblicazione: (2025)
UIClip: A Data-driven Model for Assessing User Interface Design
di: Wu, Jason, et al.
Pubblicazione: (2024)
di: Wu, Jason, et al.
Pubblicazione: (2024)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
di: Verma, Arnav, et al.
Pubblicazione: (2025)
di: Verma, Arnav, et al.
Pubblicazione: (2025)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
di: Kondic, Jovana, et al.
Pubblicazione: (2025)
di: Kondic, Jovana, et al.
Pubblicazione: (2025)
Steering Generative Models for Accessibility: EasyRead Image Generation
di: Dickenmann, Nicolas, et al.
Pubblicazione: (2026)
di: Dickenmann, Nicolas, et al.
Pubblicazione: (2026)
Computer-Use Agents as Judges for Generative User Interface
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
di: You, Keen, et al.
Pubblicazione: (2024)
di: You, Keen, et al.
Pubblicazione: (2024)
Long-Term Ad Memorability: Understanding & Generating Memorable Ads
di: SI, Harini, et al.
Pubblicazione: (2023)
di: SI, Harini, et al.
Pubblicazione: (2023)
ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots
di: Hsiao, Yu-Chung, et al.
Pubblicazione: (2022)
di: Hsiao, Yu-Chung, et al.
Pubblicazione: (2022)
Investigating Disability Representations in Text-to-Image Models
di: Tian, Yang, et al.
Pubblicazione: (2026)
di: Tian, Yang, et al.
Pubblicazione: (2026)
Can ChatGPT Read Who You Are?
di: Derner, Erik, et al.
Pubblicazione: (2023)
di: Derner, Erik, et al.
Pubblicazione: (2023)
Morae: Proactively Pausing UI Agents for User Choices
di: Peng, Yi-Hao, et al.
Pubblicazione: (2025)
di: Peng, Yi-Hao, et al.
Pubblicazione: (2025)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
di: Liu, Xinyi, et al.
Pubblicazione: (2025)
Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications
di: Asseri, Bushra, et al.
Pubblicazione: (2025)
di: Asseri, Bushra, et al.
Pubblicazione: (2025)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
di: Wang, Siting, et al.
Pubblicazione: (2025)
di: Wang, Siting, et al.
Pubblicazione: (2025)
What Color Scheme is More Effective in Assisting Readers to Locate Information in a Color-Coded Article?
di: Ng, Ho Yin, et al.
Pubblicazione: (2024)
di: Ng, Ho Yin, et al.
Pubblicazione: (2024)
True (VIS) Lies: Analyzing How Generative AI Recognizes Intentionality, Rhetoric, and Misleadingness in Visualization Lies
di: Blasilli, Graziano, et al.
Pubblicazione: (2026)
di: Blasilli, Graziano, et al.
Pubblicazione: (2026)
Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting
di: Kayser, Maxime, et al.
Pubblicazione: (2024)
di: Kayser, Maxime, et al.
Pubblicazione: (2024)
What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric
di: Kerkouri, Mohamed Amine, et al.
Pubblicazione: (2026)
di: Kerkouri, Mohamed Amine, et al.
Pubblicazione: (2026)
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
di: Mazumdar, Amrita, et al.
Pubblicazione: (2026)
di: Mazumdar, Amrita, et al.
Pubblicazione: (2026)
EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution
di: Wang, Tianfu, et al.
Pubblicazione: (2026)
di: Wang, Tianfu, et al.
Pubblicazione: (2026)
Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2024)
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2024)
Rehabilitation Exercise Quality Assessment and Feedback Generation Using Large Language Models with Prompt Engineering
di: Tang, Jessica, et al.
Pubblicazione: (2025)
di: Tang, Jessica, et al.
Pubblicazione: (2025)
How Can Large Language Models Enable Better Socially Assistive Human-Robot Interaction: A Brief Survey
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception
di: Shahzad, Sahibzada Adil, et al.
Pubblicazione: (2024)
di: Shahzad, Sahibzada Adil, et al.
Pubblicazione: (2024)
SymbolSight: Minimizing Inter-Symbol Interference for Reading with Prosthetic Vision
di: Lesner, Jasmine, et al.
Pubblicazione: (2026)
di: Lesner, Jasmine, et al.
Pubblicazione: (2026)
PromptIQ: Who Cares About Prompts? Let System Handle It -- A Component-Aware Framework for T2I Generation
di: Chhetri, Nisan, et al.
Pubblicazione: (2025)
di: Chhetri, Nisan, et al.
Pubblicazione: (2025)
Do Object Detection Localization Errors Affect Human Performance and Trust?
di: de Witte, Sven, et al.
Pubblicazione: (2024)
di: de Witte, Sven, et al.
Pubblicazione: (2024)
Documenti analoghi
-
GesGPT: Speech Gesture Synthesis With Text Parsing from ChatGPT
di: Gao, Nan, et al.
Pubblicazione: (2023) -
ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
di: Xia, Ding, et al.
Pubblicazione: (2025) -
Human-Centred Evaluation of Text-to-Image Generation Models for Self-expression of Mental Distress: A Dataset Based on GPT-4o
di: He, Sui, et al.
Pubblicazione: (2025) -
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
di: Wang, Hao, et al.
Pubblicazione: (2024) -
Can ChatGPT assist visually impaired people with micro-navigation?
di: He, Junxian, et al.
Pubblicazione: (2024)