Evolutionary Feature-wise Thresholding for Binary Representation of NLP Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sinha, Soumen, Rahnamayan, Shahryar, Bidgoli, Azam Asilian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916858375962624
author Sinha, Soumen
Rahnamayan, Shahryar
Bidgoli, Azam Asilian
author_facet Sinha, Soumen
Rahnamayan, Shahryar
Bidgoli, Azam Asilian
contents Efficient text embedding is crucial for large-scale natural language processing (NLP) applications, where storage and computational efficiency are key concerns. In this paper, we explore how using binary representations (barcodes) instead of real-valued features can be used for NLP embeddings derived from machine learning models such as BERT. Thresholding is a common method for converting continuous embeddings into binary representations, often using a fixed threshold across all features. We propose a Coordinate Search-based optimization framework that instead identifies the optimal threshold for each feature, demonstrating that feature-specific thresholds lead to improved performance in binary encoding. This ensures that the binary representations are both accurate and efficient, enhancing performance across various features. Our optimal barcode representations have shown promising results in various NLP applications, demonstrating their potential to transform text representation. We conducted extensive experiments and statistical tests on different NLP tasks and datasets to evaluate our approach and compare it to other thresholding methods. Binary embeddings generated using using optimal thresholds found by our method outperform traditional binarization methods in accuracy. This technique for generating binary representations is versatile and can be applied to any features, not just limited to NLP embeddings, making it useful for a wide range of domains in machine learning applications.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evolutionary Feature-wise Thresholding for Binary Representation of NLP Embeddings
Sinha, Soumen
Rahnamayan, Shahryar
Bidgoli, Azam Asilian
Computation and Language
Artificial Intelligence
Efficient text embedding is crucial for large-scale natural language processing (NLP) applications, where storage and computational efficiency are key concerns. In this paper, we explore how using binary representations (barcodes) instead of real-valued features can be used for NLP embeddings derived from machine learning models such as BERT. Thresholding is a common method for converting continuous embeddings into binary representations, often using a fixed threshold across all features. We propose a Coordinate Search-based optimization framework that instead identifies the optimal threshold for each feature, demonstrating that feature-specific thresholds lead to improved performance in binary encoding. This ensures that the binary representations are both accurate and efficient, enhancing performance across various features. Our optimal barcode representations have shown promising results in various NLP applications, demonstrating their potential to transform text representation. We conducted extensive experiments and statistical tests on different NLP tasks and datasets to evaluate our approach and compare it to other thresholding methods. Binary embeddings generated using using optimal thresholds found by our method outperform traditional binarization methods in accuracy. This technique for generating binary representations is versatile and can be applied to any features, not just limited to NLP embeddings, making it useful for a wide range of domains in machine learning applications.
title Evolutionary Feature-wise Thresholding for Binary Representation of NLP Embeddings
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.17025