SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Haoyu, Huang, Jinfa, Wan, Zhongwei, Zheng, Xiawu, Ji, Rongrong, Luo, Jiebo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914419005456384
author Huang, Haoyu
Huang, Jinfa
Wan, Zhongwei
Zheng, Xiawu
Ji, Rongrong
Luo, Jiebo
author_facet Huang, Haoyu
Huang, Jinfa
Wan, Zhongwei
Zheng, Xiawu
Ji, Rongrong
Luo, Jiebo
contents Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through iterative visual tool invocation. However, the cascaded perception, reasoning, and tool-calling loops introduce significant sequential overhead. This overhead, termed agentic depth, incurs prohibitive latency and seriously limits system-level concurrency. To this end, we propose SpecEyes, an agentic-level speculative acceleration framework that breaks this sequential bottleneck. Our key insight is that a lightweight, tool-free MLLM can serve as a speculative planner to predict the execution trajectory, enabling early termination of expensive tool chains without sacrificing accuracy. To regulate this speculative planning, we introduce a cognitive gating mechanism based on answer separability, which quantifies the model's confidence for self-verification without requiring oracle labels. Furthermore, we design a heterogeneous parallel funnel that exploits the stateless concurrency of the small model to mask the stateful serial execution of the large model, maximizing system throughput. Extensive experiments on V* Bench, HR-Bench, and POPE demonstrate that SpecEyes achieves 1.1-3.35x speedup over the agentic baseline while preserving or even improving accuracy (up to +6.7%), thereby boosting serving throughput under concurrent workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23483
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
Huang, Haoyu
Huang, Jinfa
Wan, Zhongwei
Zheng, Xiawu
Ji, Rongrong
Luo, Jiebo
Computer Vision and Pattern Recognition
Computation and Language
Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through iterative visual tool invocation. However, the cascaded perception, reasoning, and tool-calling loops introduce significant sequential overhead. This overhead, termed agentic depth, incurs prohibitive latency and seriously limits system-level concurrency. To this end, we propose SpecEyes, an agentic-level speculative acceleration framework that breaks this sequential bottleneck. Our key insight is that a lightweight, tool-free MLLM can serve as a speculative planner to predict the execution trajectory, enabling early termination of expensive tool chains without sacrificing accuracy. To regulate this speculative planning, we introduce a cognitive gating mechanism based on answer separability, which quantifies the model's confidence for self-verification without requiring oracle labels. Furthermore, we design a heterogeneous parallel funnel that exploits the stateless concurrency of the small model to mask the stateful serial execution of the large model, maximizing system throughput. Extensive experiments on V* Bench, HR-Bench, and POPE demonstrate that SpecEyes achieves 1.1-3.35x speedup over the agentic baseline while preserving or even improving accuracy (up to +6.7%), thereby boosting serving throughput under concurrent workloads.
title SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2603.23483