Towards Universal Semantics With Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baartmans, Raymond, Raffel, Matthew, Vikram, Rahul, Deringer, Aiden, Chen, Lizhong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915371119804416
author Baartmans, Raymond
Raffel, Matthew
Vikram, Rahul
Deringer, Aiden
Chen, Lizhong
author_facet Baartmans, Raymond
Raffel, Matthew
Vikram, Rahul
Deringer, Aiden
Chen, Lizhong
contents The Natural Semantic Metalanguage (NSM) is a linguistic theory based on a universal set of semantic primes: simple, primitive word-meanings that have been shown to exist in most, if not all, languages of the world. According to this framework, any word, regardless of complexity, can be paraphrased using these primes, revealing a clear and universally translatable meaning. These paraphrases, known as explications, can offer valuable applications for many natural language processing (NLP) tasks, but producing them has traditionally been a slow, manual process. In this work, we present the first study of using large language models (LLMs) to generate NSM explications. We introduce automatic evaluation methods, a tailored dataset for training and evaluation, and fine-tuned models for this task. Our 1B and 8B models outperform GPT-4o in producing accurate, cross-translatable explications, marking a significant step toward universal semantic representation with LLMs and opening up new possibilities for applications in semantic analysis, translation, and beyond. Our code is available at https://github.com/OSU-STARLAB/DeepNSM.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11764
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Universal Semantics With Large Language Models
Baartmans, Raymond
Raffel, Matthew
Vikram, Rahul
Deringer, Aiden
Chen, Lizhong
Computation and Language
Artificial Intelligence
The Natural Semantic Metalanguage (NSM) is a linguistic theory based on a universal set of semantic primes: simple, primitive word-meanings that have been shown to exist in most, if not all, languages of the world. According to this framework, any word, regardless of complexity, can be paraphrased using these primes, revealing a clear and universally translatable meaning. These paraphrases, known as explications, can offer valuable applications for many natural language processing (NLP) tasks, but producing them has traditionally been a slow, manual process. In this work, we present the first study of using large language models (LLMs) to generate NSM explications. We introduce automatic evaluation methods, a tailored dataset for training and evaluation, and fine-tuned models for this task. Our 1B and 8B models outperform GPT-4o in producing accurate, cross-translatable explications, marking a significant step toward universal semantic representation with LLMs and opening up new possibilities for applications in semantic analysis, translation, and beyond. Our code is available at https://github.com/OSU-STARLAB/DeepNSM.
title Towards Universal Semantics With Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.11764