A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays
Fuente:
arXiv
Saved in:
| Main Authors: | Akarajaradwong, Pawitsapak, Lertprasertphakorn, Wuttikrai, Chaksangchaichot, Chompakorn, Nutanong, Sarana |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NitiBench: A Comprehensive Study of LLM Framework Capabilities for Thai Legal Question Answering
by: Akarajaradwong, Pawitsapak, et al.
Published: (2025)
by: Akarajaradwong, Pawitsapak, et al.
Published: (2025)
Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?
by: Akarajaradwong, Pawitsapak, et al.
Published: (2025)
by: Akarajaradwong, Pawitsapak, et al.
Published: (2025)
WangchanLion and WangchanX MRC Eval
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2024)
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2024)
Evaluating Perspectival Biases in Cross-Modal Retrieval
by: Saengsukhiran, Teerapol, et al.
Published: (2025)
by: Saengsukhiran, Teerapol, et al.
Published: (2025)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
Examining the Behavior of LLM Architectures Within the Framework of Standardized National Exams in Brazil
by: Locatelli, Marcelo Sartori, et al.
Published: (2024)
by: Locatelli, Marcelo Sartori, et al.
Published: (2024)
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
Addressing Topic Leakage in Cross-Topic Evaluation for Authorship Verification
by: Sawatphol, Jitkapat, et al.
Published: (2024)
by: Sawatphol, Jitkapat, et al.
Published: (2024)
WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
by: Jiang, Han, et al.
Published: (2024)
by: Jiang, Han, et al.
Published: (2024)
Space Decomposition for Sentence Embedding
by: Ponwitayarat, Wuttikorn, et al.
Published: (2024)
by: Ponwitayarat, Wuttikorn, et al.
Published: (2024)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
by: Cao, Zhihan, et al.
Published: (2025)
by: Cao, Zhihan, et al.
Published: (2025)
Prior Prompt Engineering for Reinforcement Fine-Tuning
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
by: Tuchinda, Pume, et al.
Published: (2025)
by: Tuchinda, Pume, et al.
Published: (2025)
Examining Identity Drift in Conversations of LLM Agents
by: Choi, Junhyuk, et al.
Published: (2024)
by: Choi, Junhyuk, et al.
Published: (2024)
A Llama walks into the 'Bar': Efficient Supervised Fine-Tuning for Legal Reasoning in the Multi-state Bar Exam
by: Fernandes, Rean, et al.
Published: (2025)
by: Fernandes, Rean, et al.
Published: (2025)
Exploring Cross-Client Memorization of Training Data in Large Language Models for Federated Learning
by: Udsa, Tinnakit, et al.
Published: (2025)
by: Udsa, Tinnakit, et al.
Published: (2025)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Self-Verification is All You Need To Pass The Japanese Bar Examination
by: Shin, Andrew
Published: (2026)
by: Shin, Andrew
Published: (2026)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
by: Cai, Yunna, et al.
Published: (2025)
by: Cai, Yunna, et al.
Published: (2025)
Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
by: Sriratanawilai, Sukrit, et al.
Published: (2025)
by: Sriratanawilai, Sukrit, et al.
Published: (2025)
Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals
by: Ehara, Yo
Published: (2026)
by: Ehara, Yo
Published: (2026)
RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)
by: Cuclea, Luca-Ncolae, et al.
Published: (2026)
by: Cuclea, Luca-Ncolae, et al.
Published: (2026)
A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework
by: Li, Chenyu, et al.
Published: (2026)
by: Li, Chenyu, et al.
Published: (2026)
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
by: Chakravarty, Abhirup
Published: (2025)
by: Chakravarty, Abhirup
Published: (2025)
Emergent Social Dynamics of LLM Agents in the El Farol Bar Problem
by: Takata, Ryosuke, et al.
Published: (2025)
by: Takata, Ryosuke, et al.
Published: (2025)
Are Large Language Models Good Essay Graders?
by: Kundu, Anindita, et al.
Published: (2024)
by: Kundu, Anindita, et al.
Published: (2024)
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
by: Caraeni, Adriana, et al.
Published: (2024)
by: Caraeni, Adriana, et al.
Published: (2024)
GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations
by: Chlapanis, Odysseas S., et al.
Published: (2025)
by: Chlapanis, Odysseas S., et al.
Published: (2025)
Is GPT-4 Alone Sufficient for Automated Essay Scoring?: A Comparative Judgment Approach Based on Rater Cognition
by: Kim, Seungju, et al.
Published: (2024)
by: Kim, Seungju, et al.
Published: (2024)
Operationalizing Automated Essay Scoring: A Human-Aware Approach
by: Plasencia-Calaña, Yenisel
Published: (2025)
by: Plasencia-Calaña, Yenisel
Published: (2025)
Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
by: Shrivastava, Aryan, et al.
Published: (2024)
by: Shrivastava, Aryan, et al.
Published: (2024)
JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew
by: Razumenko, Itay, et al.
Published: (2026)
by: Razumenko, Itay, et al.
Published: (2026)
Big Bang, Low Bar -- Risk Assessment in the Public Arena
by: Price, Huw
Published: (2023)
by: Price, Huw
Published: (2023)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
by: Zhang, Bingquan, et al.
Published: (2025)
by: Zhang, Bingquan, et al.
Published: (2025)
PANORAMA: A Dataset and Benchmarks Capturing Decision Trails and Rationales in Patent Examination
by: Lim, Hyunseung, et al.
Published: (2025)
by: Lim, Hyunseung, et al.
Published: (2025)
Big5PersonalityEssays: Introducing a Novel Synthetic Generated Dataset Consisting of Short State-of-Consciousness Essays Annotated Based on the Five Factor Model of Personality
by: Floroiu, Iustin
Published: (2024)
by: Floroiu, Iustin
Published: (2024)
Language-based Valence and Arousal Expressions between the United States and China: a Cross-Cultural Examination
by: Cho, Young-Min, et al.
Published: (2024)
by: Cho, Young-Min, et al.
Published: (2024)
Similar Items
-
NitiBench: A Comprehensive Study of LLM Framework Capabilities for Thai Legal Question Answering
by: Akarajaradwong, Pawitsapak, et al.
Published: (2025) -
Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?
by: Akarajaradwong, Pawitsapak, et al.
Published: (2025) -
WangchanLion and WangchanX MRC Eval
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2024) -
Evaluating Perspectival Biases in Cross-Modal Retrieval
by: Saengsukhiran, Teerapol, et al.
Published: (2025) -
Mangosteen: An Open Thai Corpus for Language Model Pretraining
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)