Efficient Open-Vocabulary Visual Perception on Edge Devices via Multimodal Prompting

Fuente: Zenodo
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rathinasamy, Muthusami, kandasamy, saritha
Format: Recurso digital
Langue:anglais
Publié: Zenodo 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866901112360009728
author Rathinasamy, Muthusami
kandasamy, saritha
author_facet Rathinasamy, Muthusami
kandasamy, saritha
contents <p dir="auto"><strong>YOLOE-Unified</strong> is a novel framework that integrates YOLOE with distilled CLIP, runtime SAM refinement, and TensorRT optimization for efficient open-vocabulary object detection and instance segmentation on edge devices (Jetson Orin, etc.).</p> <div dir="auto"> <h3>Highlights</h3> <a href="https://github.com/muthusamir/YOLOE-Unified/blob/main/README.md#highlights"></a></div> <ul> <li>State-of-the-art zero-shot performance: <strong>48.6% mAP</strong> (detection), <strong>42.8% mask AP</strong> on LVIS</li> <li>Real-time on edge: <strong>142 FPS (FP16)</strong>, <strong>118 FPS (INT8)</strong> on Jetson Orin NX</li> <li>Low power: ~14W average consumption</li> <li>Supports text, visual, and prompt-free modes</li> </ul>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18195432
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Efficient Open-Vocabulary Visual Perception on Edge Devices via Multimodal Prompting
Rathinasamy, Muthusami
kandasamy, saritha
Visual Computing
Open-vocabulary object detection
Multimodal prompting
<p dir="auto"><strong>YOLOE-Unified</strong> is a novel framework that integrates YOLOE with distilled CLIP, runtime SAM refinement, and TensorRT optimization for efficient open-vocabulary object detection and instance segmentation on edge devices (Jetson Orin, etc.).</p> <div dir="auto"> <h3>Highlights</h3> <a href="https://github.com/muthusamir/YOLOE-Unified/blob/main/README.md#highlights"></a></div> <ul> <li>State-of-the-art zero-shot performance: <strong>48.6% mAP</strong> (detection), <strong>42.8% mask AP</strong> on LVIS</li> <li>Real-time on edge: <strong>142 FPS (FP16)</strong>, <strong>118 FPS (INT8)</strong> on Jetson Orin NX</li> <li>Low power: ~14W average consumption</li> <li>Supports text, visual, and prompt-free modes</li> </ul>
title Efficient Open-Vocabulary Visual Perception on Edge Devices via Multimodal Prompting
topic Visual Computing
Open-vocabulary object detection
Multimodal prompting
url https://doi.org/10.5281/zenodo.18195432