LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
Fuente:
arXiv
Saved in:
| Main Authors: | Son, Guijin, Ko, Hyunwoo, Lee, Hoyoung, Kim, Yewon, Hong, Seunghyeok |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
by: Son, Guijin, et al.
Published: (2026)
by: Son, Guijin, et al.
Published: (2026)
Multi-Step Reasoning in Korean and the Emergent Mirage
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
by: Ko, Hyunwoo, et al.
Published: (2025)
by: Ko, Hyunwoo, et al.
Published: (2025)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
KAIO: A Collection of More Challenging Korean Questions
by: Lee, Nahyun, et al.
Published: (2025)
by: Lee, Nahyun, et al.
Published: (2025)
Controlling Language Confusion in Multilingual LLMs
by: Lee, Nahyun, et al.
Published: (2025)
by: Lee, Nahyun, et al.
Published: (2025)
Won: Establishing Best Practices for Korean Financial NLP
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
by: Choi, Dasol, et al.
Published: (2024)
by: Choi, Dasol, et al.
Published: (2024)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
by: Choi, Dasol, et al.
Published: (2026)
by: Choi, Dasol, et al.
Published: (2026)
ResearchMath-14K: Scaling Research-Level Mathematics via Agents
by: Son, Guijin, et al.
Published: (2026)
by: Son, Guijin, et al.
Published: (2026)
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
by: Lee, Hanwool, et al.
Published: (2025)
by: Lee, Hanwool, et al.
Published: (2025)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
Revisiting the UID Hypothesis in LLM Reasoning Traces
by: Gwak, Minju, et al.
Published: (2025)
by: Gwak, Minju, et al.
Published: (2025)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
by: Gwak, Minju, et al.
Published: (2025)
by: Gwak, Minju, et al.
Published: (2025)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
by: Hong, Seokhee, et al.
Published: (2025)
by: Hong, Seokhee, et al.
Published: (2025)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
by: Kidder, William, et al.
Published: (2024)
by: Kidder, William, et al.
Published: (2024)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
On the Robustness of Reward Models for Language Model Alignment
by: Hong, Jiwoo, et al.
Published: (2025)
by: Hong, Jiwoo, et al.
Published: (2025)
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
TWICE: What Advantages Can Low-Resource Domain-Specific Embedding Model Bring? -- A Case Study on Korea Financial Texts
by: Hwang, Yewon, et al.
Published: (2025)
by: Hwang, Yewon, et al.
Published: (2025)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
by: Choi, Dasol, et al.
Published: (2025)
by: Choi, Dasol, et al.
Published: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Can Large Language Models Predict Antimicrobial Resistance Gene?
by: Yoo, Hyunwoo
Published: (2025)
by: Yoo, Hyunwoo
Published: (2025)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
by: Kumar, Divake, et al.
Published: (2026)
by: Kumar, Divake, et al.
Published: (2026)
Vision Language Models Cannot Plan, but Can They Formalize?
by: He, Muyu, et al.
Published: (2025)
by: He, Muyu, et al.
Published: (2025)
ESG Classification by Implicit Rule Learning via GPT-4
by: Yun, Hyo Jeong, et al.
Published: (2024)
by: Yun, Hyo Jeong, et al.
Published: (2024)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
by: Son, Guijin, et al.
Published: (2023)
by: Son, Guijin, et al.
Published: (2023)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
by: Schroeder, Kayla, et al.
Published: (2024)
by: Schroeder, Kayla, et al.
Published: (2024)
Can Language Models Laugh at YouTube Short-form Videos?
by: Ko, Dayoon, et al.
Published: (2023)
by: Ko, Dayoon, et al.
Published: (2023)
How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent
by: Jung, Sungwoo, et al.
Published: (2026)
by: Jung, Sungwoo, et al.
Published: (2026)
See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses
by: Chen, Yulong, et al.
Published: (2024)
by: Chen, Yulong, et al.
Published: (2024)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
by: Wu, Tianhao, et al.
Published: (2024)
by: Wu, Tianhao, et al.
Published: (2024)
AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
by: Miao, Yongliang, et al.
Published: (2026)
by: Miao, Yongliang, et al.
Published: (2026)
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
by: Chen, Jiamin, et al.
Published: (2026)
by: Chen, Jiamin, et al.
Published: (2026)
Similar Items
-
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
by: Son, Guijin, et al.
Published: (2026) -
Multi-Step Reasoning in Korean and the Emergent Mirage
by: Son, Guijin, et al.
Published: (2025) -
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
by: Ko, Hyunwoo, et al.
Published: (2025) -
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
by: Son, Guijin, et al.
Published: (2025) -
KAIO: A Collection of More Challenging Korean Questions
by: Lee, Nahyun, et al.
Published: (2025)