Hashing-Baseline: Rethinking Hashing in the Age of Pretrained Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915759366602752 |
|---|---|
| author | Moummad, Ilyass Zaher, Kawtar Rauch, Lukas Joly, Alexis |
| author_facet | Moummad, Ilyass Zaher, Kawtar Rauch, Lukas Joly, Alexis |
| contents | Information retrieval with compact binary embeddings, also referred to as hashing, is crucial for scalable fast search applications, yet state-of-the-art hashing methods require expensive, scenario-specific training. In this work, we introduce Hashing-Baseline, a strong training-free hashing method leveraging powerful pretrained encoders that produce rich pretrained embeddings. We revisit classical, training-free hashing techniques: principal component analysis, random orthogonal projection, and threshold binarization, to produce a strong baseline for hashing. Our approach combines these techniques with frozen embeddings from state-of-the-art vision and audio encoders to yield competitive retrieval performance without any additional learning or fine-tuning. To demonstrate the generality and effectiveness of this approach, we evaluate it on standard image retrieval benchmarks as well as a newly introduced benchmark for audio hashing. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_14427 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Hashing-Baseline: Rethinking Hashing in the Age of Pretrained Models Moummad, Ilyass Zaher, Kawtar Rauch, Lukas Joly, Alexis Machine Learning Information Retrieval Information retrieval with compact binary embeddings, also referred to as hashing, is crucial for scalable fast search applications, yet state-of-the-art hashing methods require expensive, scenario-specific training. In this work, we introduce Hashing-Baseline, a strong training-free hashing method leveraging powerful pretrained encoders that produce rich pretrained embeddings. We revisit classical, training-free hashing techniques: principal component analysis, random orthogonal projection, and threshold binarization, to produce a strong baseline for hashing. Our approach combines these techniques with frozen embeddings from state-of-the-art vision and audio encoders to yield competitive retrieval performance without any additional learning or fine-tuning. To demonstrate the generality and effectiveness of this approach, we evaluate it on standard image retrieval benchmarks as well as a newly introduced benchmark for audio hashing. |
| title | Hashing-Baseline: Rethinking Hashing in the Age of Pretrained Models |
| topic | Machine Learning Information Retrieval |
| url | https://arxiv.org/abs/2509.14427 |