Efficient Open-Vocabulary Visual Perception on Edge Devices via Multimodal Prompting
Fuente:
Zenodo
Enregistré dans:
| Auteurs principaux: | , |
|---|---|
| Format: | Recurso digital |
| Langue: | anglais |
| Publié: |
Zenodo
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866901112360009728 |
|---|---|
| author | Rathinasamy, Muthusami kandasamy, saritha |
| author_facet | Rathinasamy, Muthusami kandasamy, saritha |
| contents | <p dir="auto"><strong>YOLOE-Unified</strong> is a novel framework that integrates YOLOE with distilled CLIP, runtime SAM refinement, and TensorRT optimization for efficient open-vocabulary object detection and instance segmentation on edge devices (Jetson Orin, etc.).</p> <div dir="auto"> <h3>Highlights</h3> <a href="https://github.com/muthusamir/YOLOE-Unified/blob/main/README.md#highlights"></a></div> <ul> <li>State-of-the-art zero-shot performance: <strong>48.6% mAP</strong> (detection), <strong>42.8% mask AP</strong> on LVIS</li> <li>Real-time on edge: <strong>142 FPS (FP16)</strong>, <strong>118 FPS (INT8)</strong> on Jetson Orin NX</li> <li>Low power: ~14W average consumption</li> <li>Supports text, visual, and prompt-free modes</li> </ul> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18195432 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Efficient Open-Vocabulary Visual Perception on Edge Devices via Multimodal Prompting Rathinasamy, Muthusami kandasamy, saritha Visual Computing Open-vocabulary object detection Multimodal prompting <p dir="auto"><strong>YOLOE-Unified</strong> is a novel framework that integrates YOLOE with distilled CLIP, runtime SAM refinement, and TensorRT optimization for efficient open-vocabulary object detection and instance segmentation on edge devices (Jetson Orin, etc.).</p> <div dir="auto"> <h3>Highlights</h3> <a href="https://github.com/muthusamir/YOLOE-Unified/blob/main/README.md#highlights"></a></div> <ul> <li>State-of-the-art zero-shot performance: <strong>48.6% mAP</strong> (detection), <strong>42.8% mask AP</strong> on LVIS</li> <li>Real-time on edge: <strong>142 FPS (FP16)</strong>, <strong>118 FPS (INT8)</strong> on Jetson Orin NX</li> <li>Low power: ~14W average consumption</li> <li>Supports text, visual, and prompt-free modes</li> </ul> |
| title | Efficient Open-Vocabulary Visual Perception on Edge Devices via Multimodal Prompting |
| topic | Visual Computing Open-vocabulary object detection Multimodal prompting |
| url | https://doi.org/10.5281/zenodo.18195432 |