From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Stephan, Andreas, Zhu, Dawei, Aßenmacher, Matthias, Shen, Xiaoyu, Roth, Benjamin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
por: Li, Dawei, et al.
Publicado: (2024)
por: Li, Dawei, et al.
Publicado: (2024)
Counterfactual Reasoning with Knowledge Graph Embeddings
por: Zellinger, Lena, et al.
Publicado: (2024)
por: Zellinger, Lena, et al.
Publicado: (2024)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
por: Li, Dawei, et al.
Publicado: (2025)
por: Li, Dawei, et al.
Publicado: (2025)
Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
por: Wang, Zilong, et al.
Publicado: (2025)
por: Wang, Zilong, et al.
Publicado: (2025)
When Wording Steers the Evaluation: Framing Bias in LLM judges
por: Hwang, Yerin, et al.
Publicado: (2026)
por: Hwang, Yerin, et al.
Publicado: (2026)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
por: Atasoy, I. F., et al.
Publicado: (2026)
por: Atasoy, I. F., et al.
Publicado: (2026)
Examining False Positives under Inference Scaling for Mathematical Reasoning
por: Wang, Yu, et al.
Publicado: (2025)
por: Wang, Yu, et al.
Publicado: (2025)
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task
por: Yan, Yuchen, et al.
Publicado: (2025)
por: Yan, Yuchen, et al.
Publicado: (2025)
LLM-as-a-qualitative-judge: automating error analysis in natural language generation
por: Chirkova, Nadezhda, et al.
Publicado: (2025)
por: Chirkova, Nadezhda, et al.
Publicado: (2025)
Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan
por: Chatterjee, Ahan, et al.
Publicado: (2026)
por: Chatterjee, Ahan, et al.
Publicado: (2026)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
por: Cao, Hongliu, et al.
Publicado: (2025)
por: Cao, Hongliu, et al.
Publicado: (2025)
Automating Adjudication of Cardiovascular Events Using Large Language Models
por: Sivarajkumar, Sonish, et al.
Publicado: (2025)
por: Sivarajkumar, Sonish, et al.
Publicado: (2025)
LLM-Driven Multi-Turn Task-Oriented Dialogue Synthesis for Realistic Reasoning
por: Zhu, Yu, et al.
Publicado: (2026)
por: Zhu, Yu, et al.
Publicado: (2026)
No Free Lunch in Active Learning: LLM Embedding Quality Dictates Query Strategy Success
por: Rauch, Lukas, et al.
Publicado: (2025)
por: Rauch, Lukas, et al.
Publicado: (2025)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
por: Mishra, Shubhra, et al.
Publicado: (2024)
por: Mishra, Shubhra, et al.
Publicado: (2024)
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning
por: Li, Zhen, et al.
Publicado: (2025)
por: Li, Zhen, et al.
Publicado: (2025)
Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation
por: Sun, Yirong, et al.
Publicado: (2024)
por: Sun, Yirong, et al.
Publicado: (2024)
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
por: Chen, Yanjun, et al.
Publicado: (2024)
por: Chen, Yanjun, et al.
Publicado: (2024)
Saturation-Driven Dataset Generation for LLM Mathematical Reasoning in the TPTP Ecosystem
por: Quesnel, Valentin, et al.
Publicado: (2025)
por: Quesnel, Valentin, et al.
Publicado: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
por: Huan, Maggie, et al.
Publicado: (2025)
por: Huan, Maggie, et al.
Publicado: (2025)
Distilling Mathematical Reasoning Capabilities into Small Language Models
por: Zhu, Xunyu, et al.
Publicado: (2024)
por: Zhu, Xunyu, et al.
Publicado: (2024)
MultiJustice: A Chinese Dataset for Multi-Party, Multi-Charge Legal Prediction
por: Wang, Xiao, et al.
Publicado: (2025)
por: Wang, Xiao, et al.
Publicado: (2025)
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
por: Wang, Benlu, et al.
Publicado: (2025)
por: Wang, Benlu, et al.
Publicado: (2025)
Weaker LLMs' Opinions Also Matter: Mixture of Opinions Enhances LLM's Mathematical Reasoning
por: Chen, Yanan, et al.
Publicado: (2025)
por: Chen, Yanan, et al.
Publicado: (2025)
Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring
por: Dussert, Jehanne
Publicado: (2026)
por: Dussert, Jehanne
Publicado: (2026)
From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs
por: Wang, Shaojie, et al.
Publicado: (2026)
por: Wang, Shaojie, et al.
Publicado: (2026)
Reason-Align-Respond: Aligning LLM Reasoning with Knowledge Graphs for KGQA
por: Shen, Xiangqing, et al.
Publicado: (2025)
por: Shen, Xiangqing, et al.
Publicado: (2025)
Towards Robust Mathematical Reasoning
por: Luong, Thang, et al.
Publicado: (2025)
por: Luong, Thang, et al.
Publicado: (2025)
Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model
por: Zhu, Xunyu, et al.
Publicado: (2024)
por: Zhu, Xunyu, et al.
Publicado: (2024)
MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling
por: Zhu, Yakun, et al.
Publicado: (2024)
por: Zhu, Yakun, et al.
Publicado: (2024)
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
por: Huang, Yiming, et al.
Publicado: (2024)
por: Huang, Yiming, et al.
Publicado: (2024)
SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning
por: He, Jiashu, et al.
Publicado: (2025)
por: He, Jiashu, et al.
Publicado: (2025)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
por: Lee, Gisang, et al.
Publicado: (2024)
por: Lee, Gisang, et al.
Publicado: (2024)
From Test-Taking to Test-Making: Examining LLM Authoring of Commonsense Assessment Items
por: Roemmele, Melissa, et al.
Publicado: (2024)
por: Roemmele, Melissa, et al.
Publicado: (2024)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
por: Ramesh, Shyam Sundhar, et al.
Publicado: (2026)
por: Ramesh, Shyam Sundhar, et al.
Publicado: (2026)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
por: Huang, Yuzhen, et al.
Publicado: (2025)
por: Huang, Yuzhen, et al.
Publicado: (2025)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
por: Urchs, Stefanie, et al.
Publicado: (2023)
por: Urchs, Stefanie, et al.
Publicado: (2023)
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
por: Gou, Zhibin, et al.
Publicado: (2023)
por: Gou, Zhibin, et al.
Publicado: (2023)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
por: Hong, Zijin, et al.
Publicado: (2025)
por: Hong, Zijin, et al.
Publicado: (2025)
Learning From Mistakes Makes LLM Better Reasoner
por: An, Shengnan, et al.
Publicado: (2023)
por: An, Shengnan, et al.
Publicado: (2023)
Ejemplares similares
-
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
por: Li, Dawei, et al.
Publicado: (2024) -
Counterfactual Reasoning with Knowledge Graph Embeddings
por: Zellinger, Lena, et al.
Publicado: (2024) -
Preference Leakage: A Contamination Problem in LLM-as-a-judge
por: Li, Dawei, et al.
Publicado: (2025) -
Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
por: Wang, Zilong, et al.
Publicado: (2025) -
When Wording Steers the Evaluation: Framing Bias in LLM judges
por: Hwang, Yerin, et al.
Publicado: (2026)