Hashing-Baseline: Rethinking Hashing in the Age of Pretrained Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moummad, Ilyass, Zaher, Kawtar, Rauch, Lukas, Joly, Alexis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915759366602752
author Moummad, Ilyass
Zaher, Kawtar
Rauch, Lukas
Joly, Alexis
author_facet Moummad, Ilyass
Zaher, Kawtar
Rauch, Lukas
Joly, Alexis
contents Information retrieval with compact binary embeddings, also referred to as hashing, is crucial for scalable fast search applications, yet state-of-the-art hashing methods require expensive, scenario-specific training. In this work, we introduce Hashing-Baseline, a strong training-free hashing method leveraging powerful pretrained encoders that produce rich pretrained embeddings. We revisit classical, training-free hashing techniques: principal component analysis, random orthogonal projection, and threshold binarization, to produce a strong baseline for hashing. Our approach combines these techniques with frozen embeddings from state-of-the-art vision and audio encoders to yield competitive retrieval performance without any additional learning or fine-tuning. To demonstrate the generality and effectiveness of this approach, we evaluate it on standard image retrieval benchmarks as well as a newly introduced benchmark for audio hashing.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14427
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hashing-Baseline: Rethinking Hashing in the Age of Pretrained Models
Moummad, Ilyass
Zaher, Kawtar
Rauch, Lukas
Joly, Alexis
Machine Learning
Information Retrieval
Information retrieval with compact binary embeddings, also referred to as hashing, is crucial for scalable fast search applications, yet state-of-the-art hashing methods require expensive, scenario-specific training. In this work, we introduce Hashing-Baseline, a strong training-free hashing method leveraging powerful pretrained encoders that produce rich pretrained embeddings. We revisit classical, training-free hashing techniques: principal component analysis, random orthogonal projection, and threshold binarization, to produce a strong baseline for hashing. Our approach combines these techniques with frozen embeddings from state-of-the-art vision and audio encoders to yield competitive retrieval performance without any additional learning or fine-tuning. To demonstrate the generality and effectiveness of this approach, we evaluate it on standard image retrieval benchmarks as well as a newly introduced benchmark for audio hashing.
title Hashing-Baseline: Rethinking Hashing in the Age of Pretrained Models
topic Machine Learning
Information Retrieval
url https://arxiv.org/abs/2509.14427