MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Camuffo, Elena, Barbato, Francesco, Ozay, Mete, Milani, Simone, Michieli, Umberto
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917096817950720
author Camuffo, Elena
Barbato, Francesco
Ozay, Mete
Milani, Simone
Michieli, Umberto
author_facet Camuffo, Elena
Barbato, Francesco
Ozay, Mete
Milani, Simone
Michieli, Umberto
contents Personalized object detection aims to adapt a general-purpose detector to recognize user-specific instances from only a few examples. Lightweight models often struggle in this setting due to their weak semantic priors, while large vision-language models (VLMs) offer strong object-level understanding but are too computationally demanding for real-time or on-device applications. We introduce MOCHA (Multi-modal Objects-aware Cross-arcHitecture Alignment), a distillation framework that transfers multimodal region-level knowledge from a frozen VLM teacher into a lightweight vision-only detector. MOCHA extracts fused visual and textual teacher's embeddings and uses them to guide student training through a dual-objective loss that enforces accurate local alignment and global relational consistency across regions. This process enables efficient transfer of semantics without the need for teacher modifications or textual input at inference. MOCHA consistently outperforms prior baselines across four personalized detection benchmarks under strict few-shot regimes, yielding a +10.1 average improvement, with minimal inference cost.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
Camuffo, Elena
Barbato, Francesco
Ozay, Mete
Milani, Simone
Michieli, Umberto
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Personalized object detection aims to adapt a general-purpose detector to recognize user-specific instances from only a few examples. Lightweight models often struggle in this setting due to their weak semantic priors, while large vision-language models (VLMs) offer strong object-level understanding but are too computationally demanding for real-time or on-device applications. We introduce MOCHA (Multi-modal Objects-aware Cross-arcHitecture Alignment), a distillation framework that transfers multimodal region-level knowledge from a frozen VLM teacher into a lightweight vision-only detector. MOCHA extracts fused visual and textual teacher's embeddings and uses them to guide student training through a dual-objective loss that enforces accurate local alignment and global relational consistency across regions. This process enables efficient transfer of semantics without the need for teacher modifications or textual input at inference. MOCHA consistently outperforms prior baselines across four personalized detection benchmarks under strict few-shot regimes, yielding a +10.1 average improvement, with minimal inference cost.
title MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.14001