Saved in:
Bibliographic Details
Main Authors: R V, Kavin, Goyal, Pawan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.17737
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912600977047552
author R V, Kavin
Goyal, Pawan
author_facet R V, Kavin
Goyal, Pawan
contents Standard language models employ unique, monolithic embeddings for each token, potentially limiting their ability to capture the multifaceted nature of word meanings. We investigate whether tokens can be more effectively represented through a compositional structure that accumulates diverse semantic facets. To explore this, we propose Aggregate Semantic Grouping (ASG), a novel approach leveraging Product Quantization (PQ). We apply ASG to standard transformer architectures (mBERT, XLM-R, mT5) and evaluate this representational scheme across diverse tasks (NLI, NER, QA), as well as a biomedical domain-specific benchmark (BC5CDR) using BioBERT. Our findings demonstrate that representing tokens compositionally via ASG achieves extreme compression in embedding parameters (0.4--0.5\%) while maintaining $>$95\% task performance relative to the base model, even in generative tasks and extends to both cross lingual transfer and domain-specific settings. These results validate the principle that tokens can be effectively modeled as combinations of shared semantic building blocks. ASG offers a simple yet concrete method for achieving this, showcasing how compositional representations can capture linguistic richness while enabling compact yet semantically rich models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17737
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semantics
R V, Kavin
Goyal, Pawan
Computation and Language
Standard language models employ unique, monolithic embeddings for each token, potentially limiting their ability to capture the multifaceted nature of word meanings. We investigate whether tokens can be more effectively represented through a compositional structure that accumulates diverse semantic facets. To explore this, we propose Aggregate Semantic Grouping (ASG), a novel approach leveraging Product Quantization (PQ). We apply ASG to standard transformer architectures (mBERT, XLM-R, mT5) and evaluate this representational scheme across diverse tasks (NLI, NER, QA), as well as a biomedical domain-specific benchmark (BC5CDR) using BioBERT. Our findings demonstrate that representing tokens compositionally via ASG achieves extreme compression in embedding parameters (0.4--0.5\%) while maintaining $>$95\% task performance relative to the base model, even in generative tasks and extends to both cross lingual transfer and domain-specific settings. These results validate the principle that tokens can be effectively modeled as combinations of shared semantic building blocks. ASG offers a simple yet concrete method for achieving this, showcasing how compositional representations can capture linguistic richness while enabling compact yet semantically rich models.
title Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semantics
topic Computation and Language
url https://arxiv.org/abs/2509.17737