From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Stephan, Andreas, Zhu, Dawei, Aßenmacher, Matthias, Shen, Xiaoyu, Roth, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2024)
by: Li, Dawei, et al.
Published: (2024)
Counterfactual Reasoning with Knowledge Graph Embeddings
by: Zellinger, Lena, et al.
Published: (2024)
by: Zellinger, Lena, et al.
Published: (2024)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2025)
by: Li, Dawei, et al.
Published: (2025)
Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
by: Wang, Zilong, et al.
Published: (2025)
by: Wang, Zilong, et al.
Published: (2025)
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)
by: Hwang, Yerin, et al.
Published: (2026)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
by: Atasoy, I. F., et al.
Published: (2026)
by: Atasoy, I. F., et al.
Published: (2026)
Examining False Positives under Inference Scaling for Mathematical Reasoning
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task
by: Yan, Yuchen, et al.
Published: (2025)
by: Yan, Yuchen, et al.
Published: (2025)
LLM-as-a-qualitative-judge: automating error analysis in natural language generation
by: Chirkova, Nadezhda, et al.
Published: (2025)
by: Chirkova, Nadezhda, et al.
Published: (2025)
Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan
by: Chatterjee, Ahan, et al.
Published: (2026)
by: Chatterjee, Ahan, et al.
Published: (2026)
Automating Adjudication of Cardiovascular Events Using Large Language Models
by: Sivarajkumar, Sonish, et al.
Published: (2025)
by: Sivarajkumar, Sonish, et al.
Published: (2025)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
by: Cao, Hongliu, et al.
Published: (2025)
by: Cao, Hongliu, et al.
Published: (2025)
LLM-Driven Multi-Turn Task-Oriented Dialogue Synthesis for Realistic Reasoning
by: Zhu, Yu, et al.
Published: (2026)
by: Zhu, Yu, et al.
Published: (2026)
No Free Lunch in Active Learning: LLM Embedding Quality Dictates Query Strategy Success
by: Rauch, Lukas, et al.
Published: (2025)
by: Rauch, Lukas, et al.
Published: (2025)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
by: Mishra, Shubhra, et al.
Published: (2024)
by: Mishra, Shubhra, et al.
Published: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation
by: Sun, Yirong, et al.
Published: (2024)
by: Sun, Yirong, et al.
Published: (2024)
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
by: Chen, Yanjun, et al.
Published: (2024)
by: Chen, Yanjun, et al.
Published: (2024)
Saturation-Driven Dataset Generation for LLM Mathematical Reasoning in the TPTP Ecosystem
by: Quesnel, Valentin, et al.
Published: (2025)
by: Quesnel, Valentin, et al.
Published: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
Distilling Mathematical Reasoning Capabilities into Small Language Models
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
MultiJustice: A Chinese Dataset for Multi-Party, Multi-Charge Legal Prediction
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
by: Wang, Benlu, et al.
Published: (2025)
by: Wang, Benlu, et al.
Published: (2025)
Weaker LLMs' Opinions Also Matter: Mixture of Opinions Enhances LLM's Mathematical Reasoning
by: Chen, Yanan, et al.
Published: (2025)
by: Chen, Yanan, et al.
Published: (2025)
Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring
by: Dussert, Jehanne
Published: (2026)
by: Dussert, Jehanne
Published: (2026)
From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs
by: Wang, Shaojie, et al.
Published: (2026)
by: Wang, Shaojie, et al.
Published: (2026)
Reason-Align-Respond: Aligning LLM Reasoning with Knowledge Graphs for KGQA
by: Shen, Xiangqing, et al.
Published: (2025)
by: Shen, Xiangqing, et al.
Published: (2025)
Towards Robust Mathematical Reasoning
by: Luong, Thang, et al.
Published: (2025)
by: Luong, Thang, et al.
Published: (2025)
Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
by: Huang, Yiming, et al.
Published: (2024)
by: Huang, Yiming, et al.
Published: (2024)
MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling
by: Zhu, Yakun, et al.
Published: (2024)
by: Zhu, Yakun, et al.
Published: (2024)
SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning
by: He, Jiashu, et al.
Published: (2025)
by: He, Jiashu, et al.
Published: (2025)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
by: Lee, Gisang, et al.
Published: (2024)
by: Lee, Gisang, et al.
Published: (2024)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
by: Roemmele, Melissa, et al.
Published: (2024)
by: Roemmele, Melissa, et al.
Published: (2024)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
by: Urchs, Stefanie, et al.
Published: (2023)
by: Urchs, Stefanie, et al.
Published: (2023)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
by: Huang, Yuzhen, et al.
Published: (2025)
by: Huang, Yuzhen, et al.
Published: (2025)
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
by: Gou, Zhibin, et al.
Published: (2023)
by: Gou, Zhibin, et al.
Published: (2023)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
by: Hong, Zijin, et al.
Published: (2025)
by: Hong, Zijin, et al.
Published: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
Similar Items
-
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2024) -
Counterfactual Reasoning with Knowledge Graph Embeddings
by: Zellinger, Lena, et al.
Published: (2024) -
Preference Leakage: A Contamination Problem in LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2025) -
Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
by: Wang, Zilong, et al.
Published: (2025) -
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)