NitiBench: A Comprehensive Study of LLM Framework Capabilities for Thai Legal Question Answering
Fuente:
arXiv
Salvato in:
| Autori principali: | Akarajaradwong, Pawitsapak, Pothavorn, Pirat, Chaksangchaichot, Chompakorn, Tasawong, Panuthep, Nopparatbundit, Thitiwat, Pratai, Keerakiat, Nutanong, Sarana |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2025)
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2025)
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2026)
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2026)
WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
Evaluating Perspectival Biases in Cross-Modal Retrieval
di: Saengsukhiran, Teerapol, et al.
Pubblicazione: (2025)
di: Saengsukhiran, Teerapol, et al.
Pubblicazione: (2025)
WangchanLion and WangchanX MRC Eval
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2024)
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2024)
THAI Speech Emotion Recognition (THAI-SER) corpus
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
di: Wongpithayadisai, Jilamika, et al.
Pubblicazione: (2025)
SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
di: Tasawong, Panuthep, et al.
Pubblicazione: (2025)
di: Tasawong, Panuthep, et al.
Pubblicazione: (2025)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2025)
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2025)
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
Addressing Topic Leakage in Cross-Topic Evaluation for Authorship Verification
di: Sawatphol, Jitkapat, et al.
Pubblicazione: (2024)
di: Sawatphol, Jitkapat, et al.
Pubblicazione: (2024)
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2025)
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2025)
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
di: Tasawong, Panuthep, et al.
Pubblicazione: (2026)
di: Tasawong, Panuthep, et al.
Pubblicazione: (2026)
Space Decomposition for Sentence Embedding
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2024)
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2024)
Prior Prompt Engineering for Reinforcement Fine-Tuning
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering
di: AlMannaa, Feras, et al.
Pubblicazione: (2025)
di: AlMannaa, Feras, et al.
Pubblicazione: (2025)
Probability and statistical inference / Nitis Mukhopadhyay
di: Mukhopadhyay, Nitis
di: Mukhopadhyay, Nitis
Exploring Cross-Client Memorization of Training Data in Large Language Models for Federated Learning
di: Udsa, Tinnakit, et al.
Pubblicazione: (2025)
di: Udsa, Tinnakit, et al.
Pubblicazione: (2025)
When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
di: Tuchinda, Pume, et al.
Pubblicazione: (2025)
di: Tuchinda, Pume, et al.
Pubblicazione: (2025)
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models
di: Li, Linyi, et al.
Pubblicazione: (2024)
di: Li, Linyi, et al.
Pubblicazione: (2024)
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
di: Wu, Xianjie, et al.
Pubblicazione: (2024)
di: Wu, Xianjie, et al.
Pubblicazione: (2024)
Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
di: Sriratanawilai, Sukrit, et al.
Pubblicazione: (2025)
di: Sriratanawilai, Sukrit, et al.
Pubblicazione: (2025)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering
di: Nguyen, Tan-Minh, et al.
Pubblicazione: (2025)
di: Nguyen, Tan-Minh, et al.
Pubblicazione: (2025)
Sonographic appearance of focal liver lesions and likelihood of hepatocellular carcinoma in adult Thais with chronic hepatitis B virus infection
di: Sarana Suttivanich, et al.
Pubblicazione: (2024)
di: Sarana Suttivanich, et al.
Pubblicazione: (2024)
Answer Retrieval in Legal Community Question Answering
di: Askari, Arian, et al.
Pubblicazione: (2024)
di: Askari, Arian, et al.
Pubblicazione: (2024)
Measuring the Groundedness of Legal Question-Answering Systems
di: Trautmann, Dietrich, et al.
Pubblicazione: (2024)
di: Trautmann, Dietrich, et al.
Pubblicazione: (2024)
Intelligent Legal Assistant: An Interactive Clarification System for Legal Question Answering
di: Yao, Rujing, et al.
Pubblicazione: (2025)
di: Yao, Rujing, et al.
Pubblicazione: (2025)
PAKTON: A Multi-Agent Framework for Question Answering in Long Legal Agreements
di: Raptopoulos, Petros, et al.
Pubblicazione: (2025)
di: Raptopoulos, Petros, et al.
Pubblicazione: (2025)
Vietnamese Legal Information Retrieval in Question-Answering System
di: Ba, Thiem Nguyen, et al.
Pubblicazione: (2024)
di: Ba, Thiem Nguyen, et al.
Pubblicazione: (2024)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
di: Lee, Gyubok, et al.
Pubblicazione: (2025)
di: Lee, Gyubok, et al.
Pubblicazione: (2025)
JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability
di: Wang, Junda, et al.
Pubblicazione: (2024)
di: Wang, Junda, et al.
Pubblicazione: (2024)
Long-Context Long-Form Question Answering for Legal Domain
di: Kulkarni, Anagha, et al.
Pubblicazione: (2026)
di: Kulkarni, Anagha, et al.
Pubblicazione: (2026)
Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval
di: lee, Jihyung, et al.
Pubblicazione: (2026)
di: lee, Jihyung, et al.
Pubblicazione: (2026)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
di: Li, Haitao, et al.
Pubblicazione: (2024)
di: Li, Haitao, et al.
Pubblicazione: (2024)
A Comprehensive Graph Framework for Question Answering with Mode-Seeking Preference Alignment
di: Tang, Quanwei, et al.
Pubblicazione: (2025)
di: Tang, Quanwei, et al.
Pubblicazione: (2025)
ThaiSafetyBench: Assessing Language Model Safety in Thai Cultural Contexts
di: Ukarapol, Trapoom, et al.
Pubblicazione: (2026)
di: Ukarapol, Trapoom, et al.
Pubblicazione: (2026)
Pre-training, Fine-tuning and Re-ranking: A Three-Stage Framework for Legal Question Answering
di: Ni, Shiwen, et al.
Pubblicazione: (2024)
di: Ni, Shiwen, et al.
Pubblicazione: (2024)
A Simple LLM Framework for Long-Range Video Question-Answering
di: Zhang, Ce, et al.
Pubblicazione: (2023)
di: Zhang, Ce, et al.
Pubblicazione: (2023)
Experimenting with Legal AI Solutions: The Case of Question-Answering for Access to Justice
di: Li, Jonathan, et al.
Pubblicazione: (2024)
di: Li, Jonathan, et al.
Pubblicazione: (2024)
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
di: Lee, Jihyung, et al.
Pubblicazione: (2025)
di: Lee, Jihyung, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2025) -
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2026) -
WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025) -
Evaluating Perspectival Biases in Cross-Modal Retrieval
di: Saengsukhiran, Teerapol, et al.
Pubblicazione: (2025) -
WangchanLion and WangchanX MRC Eval
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2024)