Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akarajaradwong, Pawitsapak, Chaksangchaichot, Chompakorn, Pothavorn, Pirat, Thamrongrattanarit-Rutherford, Attapol, Chuangsuwanich, Ekapol, Nutanong, Sarana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912479508955136
author Akarajaradwong, Pawitsapak
Chaksangchaichot, Chompakorn
Pothavorn, Pirat
Thamrongrattanarit-Rutherford, Attapol
Chuangsuwanich, Ekapol
Nutanong, Sarana
author_facet Akarajaradwong, Pawitsapak
Chaksangchaichot, Chompakorn
Pothavorn, Pirat
Thamrongrattanarit-Rutherford, Attapol
Chuangsuwanich, Ekapol
Nutanong, Sarana
contents The Retrieval-Augmented Generation (RAG) systems' performance on Thai legal question answering is still limited, especially for questions requiring extensive, complex legal reasoning. To address these limitations, we introduce an approach aligning LLMs toward improved law citation accuracy and better response quality using Group-Relative Policy Optimization (GRPO). Our approach leverages BGE-M3 embeddings as a cost-efficient semantic-similarity reward, significantly reducing computational expenses up to 2.5x compared to large language model judges. Experiments on the NitiBench benchmark demonstrate substantial improvements: GRPO achieves up to 90% citation-F1 gains from the base model and a 31% increase in joint quality metrics over instruction tuning. Crucially, our method shows enhanced robustness on complex legal reasoning tasks compared to instruction tuning, providing an effective and resource-efficient solution for enhancing Thai legal LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09638
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?
Akarajaradwong, Pawitsapak
Chaksangchaichot, Chompakorn
Pothavorn, Pirat
Thamrongrattanarit-Rutherford, Attapol
Chuangsuwanich, Ekapol
Nutanong, Sarana
Computation and Language
The Retrieval-Augmented Generation (RAG) systems' performance on Thai legal question answering is still limited, especially for questions requiring extensive, complex legal reasoning. To address these limitations, we introduce an approach aligning LLMs toward improved law citation accuracy and better response quality using Group-Relative Policy Optimization (GRPO). Our approach leverages BGE-M3 embeddings as a cost-efficient semantic-similarity reward, significantly reducing computational expenses up to 2.5x compared to large language model judges. Experiments on the NitiBench benchmark demonstrate substantial improvements: GRPO achieves up to 90% citation-F1 gains from the base model and a 31% increase in joint quality metrics over instruction tuning. Crucially, our method shows enhanced robustness on complex legal reasoning tasks compared to instruction tuning, providing an effective and resource-efficient solution for enhancing Thai legal LLMs.
title Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?
topic Computation and Language
url https://arxiv.org/abs/2507.09638