Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chhoun, Sovandara, Po, Pichdara, Ros, Sereiwathna, Cho, Wan-Sup, Khoeurn, Saksonita
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910245460115456
author Chhoun, Sovandara
Po, Pichdara
Ros, Sereiwathna
Cho, Wan-Sup
Khoeurn, Saksonita
author_facet Chhoun, Sovandara
Po, Pichdara
Ros, Sereiwathna
Cho, Wan-Sup
Khoeurn, Saksonita
contents In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware, Sentence-Based, and LLM-Based within a Retrieval-Augmented Generation (RAG) framework applied to Khmer agricultural documents. The document chunks are encoded using the BGE-M3 multilingual embedding model and retrieved using the FAISS library. Performance is evaluated using four metrics: Average Retrieval Score (L2 distance), Answer Relevance, Khmer Coverage, and Khmer Intersection over Union, all measured against ground-truth question-answer pairs. For evaluation, we perform 5-fold cross-validation over 18 question-answer pairs. We observe the best performance for the character-based Recursive chunking method with a chunk size of 300 characters, achieving the lowest L2 distance (0.4295 +- 0.0461), highest Answer Relevance (0.8663 +- 0.0199), and highest Khmer IoU (0.6441 +- 0.0347). A paired t-test shows a statistically significant improvement over the Sentence-Based chunking method in L2 distance (p = 0.0121). These results highlight the importance of segmentation granularity and structural preservation for optimizing dense retrieval in morphologically complex, low-resource languages such as Khmer.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22203
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents
Chhoun, Sovandara
Po, Pichdara
Ros, Sereiwathna
Cho, Wan-Sup
Khoeurn, Saksonita
Computation and Language
H.3.3; I.2.7
In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware, Sentence-Based, and LLM-Based within a Retrieval-Augmented Generation (RAG) framework applied to Khmer agricultural documents. The document chunks are encoded using the BGE-M3 multilingual embedding model and retrieved using the FAISS library. Performance is evaluated using four metrics: Average Retrieval Score (L2 distance), Answer Relevance, Khmer Coverage, and Khmer Intersection over Union, all measured against ground-truth question-answer pairs. For evaluation, we perform 5-fold cross-validation over 18 question-answer pairs. We observe the best performance for the character-based Recursive chunking method with a chunk size of 300 characters, achieving the lowest L2 distance (0.4295 +- 0.0461), highest Answer Relevance (0.8663 +- 0.0199), and highest Khmer IoU (0.6441 +- 0.0347). A paired t-test shows a statistically significant improvement over the Sentence-Based chunking method in L2 distance (p = 0.0121). These results highlight the importance of segmentation granularity and structural preservation for optimizing dense retrieval in morphologically complex, low-resource languages such as Khmer.
title Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents
topic Computation and Language
H.3.3; I.2.7
url https://arxiv.org/abs/2605.22203