SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ponwitayarat, Wuttikorn, Limkonchotiwat, Peerat, Ng, Raymond, Montalan, Jann Railey, Aung, Thura, Ngui, Jian Gang, Susanto, Yosephine, Tjhi, William Chandra, Tasawong, Panuthep, Cambria, Erik, Chuangsuwanich, Ekapol, Nutanong, Sarana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915922509299712
author Ponwitayarat, Wuttikorn
Limkonchotiwat, Peerat
Ng, Raymond
Montalan, Jann Railey
Aung, Thura
Ngui, Jian Gang
Susanto, Yosephine
Tjhi, William Chandra
Tasawong, Panuthep
Cambria, Erik
Chuangsuwanich, Ekapol
Nutanong, Sarana
author_facet Ponwitayarat, Wuttikorn
Limkonchotiwat, Peerat
Ng, Raymond
Montalan, Jann Railey
Aung, Thura
Ngui, Jian Gang
Susanto, Yosephine
Tjhi, William Chandra
Tasawong, Panuthep
Cambria, Erik
Chuangsuwanich, Ekapol
Nutanong, Sarana
contents Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our results show that this assumption does not hold in practice. We introduce SEA-BED, a large-scale benchmark covering 10 Southeast Asian (SEA) languages and diverse embedding tasks, designed to systematically examine how embedding performance varies across tasks, languages, and language-task combinations. Across extensive evaluations, we observe that no single model performs uniformly well across SEA languages; task difficulty differs markedly within languages, and success on one task does not reliably generalize to others. Language-task analyses further reveal highly non-uniform performance landscapes, where performance varies across different language-task combinations. These findings call for closer attention to performance measurements that provide an expansive view across languages and tasks to uncover inconsistencies in semantic representation. Based on these observations, we provide insights for future model development, including data, algorithmic, and architectural considerations.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12243
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
Ponwitayarat, Wuttikorn
Limkonchotiwat, Peerat
Ng, Raymond
Montalan, Jann Railey
Aung, Thura
Ngui, Jian Gang
Susanto, Yosephine
Tjhi, William Chandra
Tasawong, Panuthep
Cambria, Erik
Chuangsuwanich, Ekapol
Nutanong, Sarana
Computation and Language
Multilingual text embeddings are often assumed to encode meaning in a perspective-independent semantic space, yielding stable similarity judgments across tasks and languages. Our results show that this assumption does not hold in practice. We introduce SEA-BED, a large-scale benchmark covering 10 Southeast Asian (SEA) languages and diverse embedding tasks, designed to systematically examine how embedding performance varies across tasks, languages, and language-task combinations. Across extensive evaluations, we observe that no single model performs uniformly well across SEA languages; task difficulty differs markedly within languages, and success on one task does not reliably generalize to others. Language-task analyses further reveal highly non-uniform performance landscapes, where performance varies across different language-task combinations. These findings call for closer attention to performance measurements that provide an expansive view across languages and tasks to uncover inconsistencies in semantic representation. Based on these observations, we provide insights for future model development, including data, algorithmic, and architectural considerations.
title SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
topic Computation and Language
url https://arxiv.org/abs/2508.12243