Token-Efficient Multimodal Reasoning via Image Prompt Packaging

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Choi, Joong Ho, Zhao, Jiayang, Appalla, Avani, Mukesh, Himansh, Vasani, Dhwanil, Qian, Boyi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917382286475264
author Choi, Joong Ho
Zhao, Jiayang
Appalla, Avani
Mukesh, Himansh
Vasani, Dhwanil
Qian, Boyi
author_facet Choi, Joong Ho
Zhao, Jiayang
Appalla, Avani
Mukesh, Himansh
Vasani, Dhwanil
Qian, Boyi
contents Deploying large multimodal language models at scale is constrained by token-based inference costs, yet the cost-performance behavior of visual prompting strategies remains poorly characterized. We introduce Image Prompt Packaging (IPPg), a prompting paradigm that embeds structured text directly into images to reduce text token overhead, and benchmark it across five datasets, three frontier models (GPT-4.1, GPT-4o, Claude 3.5 Sonnet), and two task families (VQA and code generation). We derive a cost formulation decomposing savings by token type and show IPPg achieves 35.8--91.0\% inference cost reductions. Despite token compression of up to 96\%, accuracy remains competitive in many settings, though outcomes are highly model- and task-dependent: GPT-4.1 achieves simultaneous accuracy and cost gains on CoSQL, while Claude 3.5 incurs cost increases on several VQA benchmarks. Systematic error analysis yields a failure-mode taxonomy: spatial reasoning, non-English inputs, and character-sensitive operations are most vulnerable, while schema-structured tasks benefit most. A 125-configuration rendering ablation reveals accuracy shifts of 10--30 percentage points, establishing visual encoding choices as a first-class variable in multimodal system design.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02492
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Token-Efficient Multimodal Reasoning via Image Prompt Packaging
Choi, Joong Ho
Zhao, Jiayang
Appalla, Avani
Mukesh, Himansh
Vasani, Dhwanil
Qian, Boyi
Computer Vision and Pattern Recognition
Artificial Intelligence
Deploying large multimodal language models at scale is constrained by token-based inference costs, yet the cost-performance behavior of visual prompting strategies remains poorly characterized. We introduce Image Prompt Packaging (IPPg), a prompting paradigm that embeds structured text directly into images to reduce text token overhead, and benchmark it across five datasets, three frontier models (GPT-4.1, GPT-4o, Claude 3.5 Sonnet), and two task families (VQA and code generation). We derive a cost formulation decomposing savings by token type and show IPPg achieves 35.8--91.0\% inference cost reductions. Despite token compression of up to 96\%, accuracy remains competitive in many settings, though outcomes are highly model- and task-dependent: GPT-4.1 achieves simultaneous accuracy and cost gains on CoSQL, while Claude 3.5 incurs cost increases on several VQA benchmarks. Systematic error analysis yields a failure-mode taxonomy: spatial reasoning, non-English inputs, and character-sensitive operations are most vulnerable, while schema-structured tasks benefit most. A 125-configuration rendering ablation reveals accuracy shifts of 10--30 percentage points, establishing visual encoding choices as a first-class variable in multimodal system design.
title Token-Efficient Multimodal Reasoning via Image Prompt Packaging
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.02492