REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Khosla, Savya, TV, Sethuraman, Lee, Barnett, Schwing, Alexander, Hoiem, Derek
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915590868828160
author Khosla, Savya
TV, Sethuraman
Lee, Barnett
Schwing, Alexander
Hoiem, Derek
author_facet Khosla, Savya
TV, Sethuraman
Lee, Barnett
Schwing, Alexander
Hoiem, Derek
contents We introduce the Region Encoder Network (REN), a fast and effective model for generating region-based image representations using point prompts. Recent methods combine class-agnostic segmenters (e.g., SAM) with patch-based image encoders (e.g., DINO) to produce compact and effective region representations, but they suffer from high computational cost due to the segmentation step. REN bypasses this bottleneck using a lightweight module that directly generates region tokens, enabling 60x faster token generation with 35x less memory, while also improving token quality. It uses a few cross-attention blocks that take point prompts as queries and features from a patch-based image encoder as keys and values to produce region tokens that correspond to the prompted objects. We train REN with three popular encoders-DINO, DINOv2, and OpenCLIP-and show that it can be extended to other encoders without dedicated training. We evaluate REN on semantic segmentation and retrieval tasks, where it consistently outperforms the original encoders in both performance and compactness, and matches or exceeds SAM-based region methods while being significantly faster. Notably, REN achieves state-of-the-art results on the challenging Ego4D VQ2D benchmark and outperforms proprietary LMMs on Visual Haystacks' single-needle challenge. Code and models are available at: https://github.com/savya08/REN.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18153
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
Khosla, Savya
TV, Sethuraman
Lee, Barnett
Schwing, Alexander
Hoiem, Derek
Computer Vision and Pattern Recognition
We introduce the Region Encoder Network (REN), a fast and effective model for generating region-based image representations using point prompts. Recent methods combine class-agnostic segmenters (e.g., SAM) with patch-based image encoders (e.g., DINO) to produce compact and effective region representations, but they suffer from high computational cost due to the segmentation step. REN bypasses this bottleneck using a lightweight module that directly generates region tokens, enabling 60x faster token generation with 35x less memory, while also improving token quality. It uses a few cross-attention blocks that take point prompts as queries and features from a patch-based image encoder as keys and values to produce region tokens that correspond to the prompted objects. We train REN with three popular encoders-DINO, DINOv2, and OpenCLIP-and show that it can be extended to other encoders without dedicated training. We evaluate REN on semantic segmentation and retrieval tasks, where it consistently outperforms the original encoders in both performance and compactness, and matches or exceeds SAM-based region methods while being significantly faster. Notably, REN achieves state-of-the-art results on the challenging Ego4D VQ2D benchmark and outperforms proprietary LMMs on Visual Haystacks' single-needle challenge. Code and models are available at: https://github.com/savya08/REN.
title REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.18153