Discovering Semantic Subdimensions through Disentangled Conceptual Representations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Yunhao, Wang, Shaonan, Lin, Nan, Dong, Xinyi, Li, Chong, Zong, Chengqing
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909796984160256
author Zhang, Yunhao
Wang, Shaonan
Lin, Nan
Dong, Xinyi
Li, Chong
Zong, Chengqing
author_facet Zhang, Yunhao
Wang, Shaonan
Lin, Nan
Dong, Xinyi
Li, Chong
Zong, Chengqing
contents Understanding the core dimensions of conceptual semantics is fundamental to uncovering how meaning is organized in language and the brain. Existing approaches often rely on predefined semantic dimensions that offer only broad representations, overlooking finer conceptual distinctions. This paper proposes a novel framework to investigate the subdimensions underlying coarse-grained semantic dimensions. Specifically, we introduce a Disentangled Continuous Semantic Representation Model (DCSRM) that decomposes word embeddings from large language models into multiple sub-embeddings, each encoding specific semantic information. Using these sub-embeddings, we identify a set of interpretable semantic subdimensions. To assess their neural plausibility, we apply voxel-wise encoding models to map these subdimensions to brain activation. Our work offers more fine-grained interpretable semantic subdimensions of conceptual meaning. Further analyses reveal that semantic dimensions are structured according to distinct principles, with polarity emerging as a key factor driving their decomposition into subdimensions. The neural correlates of the identified subdimensions support their cognitive and neuroscientific plausibility.
format Preprint
id arxiv_https___arxiv_org_abs_2508_21436
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discovering Semantic Subdimensions through Disentangled Conceptual Representations
Zhang, Yunhao
Wang, Shaonan
Lin, Nan
Dong, Xinyi
Li, Chong
Zong, Chengqing
Computation and Language
Understanding the core dimensions of conceptual semantics is fundamental to uncovering how meaning is organized in language and the brain. Existing approaches often rely on predefined semantic dimensions that offer only broad representations, overlooking finer conceptual distinctions. This paper proposes a novel framework to investigate the subdimensions underlying coarse-grained semantic dimensions. Specifically, we introduce a Disentangled Continuous Semantic Representation Model (DCSRM) that decomposes word embeddings from large language models into multiple sub-embeddings, each encoding specific semantic information. Using these sub-embeddings, we identify a set of interpretable semantic subdimensions. To assess their neural plausibility, we apply voxel-wise encoding models to map these subdimensions to brain activation. Our work offers more fine-grained interpretable semantic subdimensions of conceptual meaning. Further analyses reveal that semantic dimensions are structured according to distinct principles, with polarity emerging as a key factor driving their decomposition into subdimensions. The neural correlates of the identified subdimensions support their cognitive and neuroscientific plausibility.
title Discovering Semantic Subdimensions through Disentangled Conceptual Representations
topic Computation and Language
url https://arxiv.org/abs/2508.21436