Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yunqing, Yang, Zheming, Zhao, Chang, Ji, Wen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908556617318400
author Hu, Yunqing
Yang, Zheming
Zhao, Chang
Ji, Wen
author_facet Hu, Yunqing
Yang, Zheming
Zhao, Chang
Ji, Wen
contents Traditional object detection methods face performance degradation challenges in complex scenarios such as low-light conditions and heavy occlusions due to a lack of high-level semantic understanding. To address this, this paper proposes an adaptive guidance-based semantic enhancement edge-cloud collaborative object detection method leveraging Multimodal Large Language Models (MLLM), achieving an effective balance between accuracy and efficiency. Specifically, the method first employs instruction fine-tuning to enable the MLLM to generate structured scene descriptions. It then designs an adaptive mapping mechanism that dynamically converts semantic information into parameter adjustment signals for edge detectors, achieving real-time semantic enhancement. Within an edge-cloud collaborative inference framework, the system automatically selects between invoking cloud-based semantic guidance or directly outputting edge detection results based on confidence scores. Experiments demonstrate that the proposed method effectively enhances detection accuracy and efficiency in complex scenes. Specifically, it can reduce latency by over 79% and computational cost by 70% in low-light and highly occluded scenes while maintaining accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19875
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
Hu, Yunqing
Yang, Zheming
Zhao, Chang
Ji, Wen
Computer Vision and Pattern Recognition
Artificial Intelligence
Traditional object detection methods face performance degradation challenges in complex scenarios such as low-light conditions and heavy occlusions due to a lack of high-level semantic understanding. To address this, this paper proposes an adaptive guidance-based semantic enhancement edge-cloud collaborative object detection method leveraging Multimodal Large Language Models (MLLM), achieving an effective balance between accuracy and efficiency. Specifically, the method first employs instruction fine-tuning to enable the MLLM to generate structured scene descriptions. It then designs an adaptive mapping mechanism that dynamically converts semantic information into parameter adjustment signals for edge detectors, achieving real-time semantic enhancement. Within an edge-cloud collaborative inference framework, the system automatically selects between invoking cloud-based semantic guidance or directly outputting edge detection results based on confidence scores. Experiments demonstrate that the proposed method effectively enhances detection accuracy and efficiency in complex scenes. Specifically, it can reduce latency by over 79% and computational cost by 70% in low-light and highly occluded scenes while maintaining accuracy.
title Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.19875