SailCompass: Towards Reproducible and Robust Evaluation for Southeast Asian Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Jia, Dou, Longxu, Zeng, Guangtao, Kok, Stanley, Lu, Wei, Liu, Qian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929610489331712
author Guo, Jia
Dou, Longxu
Zeng, Guangtao
Kok, Stanley
Lu, Wei
Liu, Qian
author_facet Guo, Jia
Dou, Longxu
Zeng, Guangtao
Kok, Stanley
Lu, Wei
Liu, Qian
contents In this paper, we introduce SailCompass, a reproducible and robust evaluation benchmark for assessing Large Language Models (LLMs) on Southeast Asian Languages (SEA). SailCompass encompasses three main SEA languages, eight primary tasks including 14 datasets covering three task types (generation, multiple-choice questions, and classification). To improve the robustness of the evaluation approach, we explore different prompt configurations for multiple-choice questions and leverage calibrations to improve the faithfulness of classification tasks. With SailCompass, we derive the following findings: (1) SEA-specialized LLMs still outperform general LLMs, although the gap has narrowed; (2) A balanced language distribution is important for developing better SEA-specialized LLMs; (3) Advanced prompting techniques (e.g., calibration, perplexity-based ranking) are necessary to better utilize LLMs. All datasets and evaluation scripts are public.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01186
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SailCompass: Towards Reproducible and Robust Evaluation for Southeast Asian Languages
Guo, Jia
Dou, Longxu
Zeng, Guangtao
Kok, Stanley
Lu, Wei
Liu, Qian
Computation and Language
In this paper, we introduce SailCompass, a reproducible and robust evaluation benchmark for assessing Large Language Models (LLMs) on Southeast Asian Languages (SEA). SailCompass encompasses three main SEA languages, eight primary tasks including 14 datasets covering three task types (generation, multiple-choice questions, and classification). To improve the robustness of the evaluation approach, we explore different prompt configurations for multiple-choice questions and leverage calibrations to improve the faithfulness of classification tasks. With SailCompass, we derive the following findings: (1) SEA-specialized LLMs still outperform general LLMs, although the gap has narrowed; (2) A balanced language distribution is important for developing better SEA-specialized LLMs; (3) Advanced prompting techniques (e.g., calibration, perplexity-based ranking) are necessary to better utilize LLMs. All datasets and evaluation scripts are public.
title SailCompass: Towards Reproducible and Robust Evaluation for Southeast Asian Languages
topic Computation and Language
url https://arxiv.org/abs/2412.01186