Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Lan, Valentino, Marco, Freitas, Andre |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions
por: Zhang, Lan, et al.
Publicado: (2025)
por: Zhang, Lan, et al.
Publicado: (2025)
MASA: LLM-Driven Multi-Agent Systems for Autoformalization
por: Zhang, Lan, et al.
Publicado: (2025)
por: Zhang, Lan, et al.
Publicado: (2025)
PEIRCE: Unifying Material and Formal Reasoning via LLM-Driven Neuro-Symbolic Refinement
por: Quan, Xin, et al.
Publicado: (2025)
por: Quan, Xin, et al.
Publicado: (2025)
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
por: Meadows, Jordan, et al.
Publicado: (2023)
por: Meadows, Jordan, et al.
Publicado: (2023)
Reasoning with Natural Language Explanations
por: Valentino, Marco, et al.
Publicado: (2024)
por: Valentino, Marco, et al.
Publicado: (2024)
Monotonic Reference-Free Refinement for Autoformalization
por: Zhang, Lan, et al.
Publicado: (2026)
por: Zhang, Lan, et al.
Publicado: (2026)
Adaptive LLM-Symbolic Reasoning via Dynamic Logical Solver Composition
por: Xu, Lei, et al.
Publicado: (2025)
por: Xu, Lei, et al.
Publicado: (2025)
Improving Chain-of-Thought Reasoning via Quasi-Symbolic Abstractions
por: Ranaldi, Leonardo, et al.
Publicado: (2025)
por: Ranaldi, Leonardo, et al.
Publicado: (2025)
Controlling Equational Reasoning in Large Language Models with Prompt Interventions
por: Meadows, Jordan, et al.
Publicado: (2023)
por: Meadows, Jordan, et al.
Publicado: (2023)
Learning to Disentangle Latent Reasoning Rules with Language VAEs: A Systematic Study
por: Zhang, Yingji, et al.
Publicado: (2025)
por: Zhang, Yingji, et al.
Publicado: (2025)
Logic-Parametric Neuro-Symbolic NLI: Controlling Logical Formalisms for Verifiable LLM Reasoning
por: Farjami, Ali, et al.
Publicado: (2026)
por: Farjami, Ali, et al.
Publicado: (2026)
Eliciting Critical Reasoning in Retrieval-Augmented Language Models via Contrastive Explanations
por: Ranaldi, Leonardo, et al.
Publicado: (2024)
por: Ranaldi, Leonardo, et al.
Publicado: (2024)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
por: Kim, Geonhee, et al.
Publicado: (2024)
por: Kim, Geonhee, et al.
Publicado: (2024)
Inferring Latent Intentions: Attributional Natural Language Inference in LLM Agents
por: Quan, Xin, et al.
Publicado: (2026)
por: Quan, Xin, et al.
Publicado: (2026)
Consistent Autoformalization for Constructing Mathematical Libraries
por: Zhang, Lan, et al.
Publicado: (2024)
por: Zhang, Lan, et al.
Publicado: (2024)
Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving
por: Quan, Xin, et al.
Publicado: (2024)
por: Quan, Xin, et al.
Publicado: (2024)
Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies
por: Jullien, Mael, et al.
Publicado: (2025)
por: Jullien, Mael, et al.
Publicado: (2025)
Estimating the Causal Effects of Natural Logic Features in Neural NLI Models
por: Rozanova, Julia, et al.
Publicado: (2023)
por: Rozanova, Julia, et al.
Publicado: (2023)
Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation
por: Aljaafari, Nura, et al.
Publicado: (2026)
por: Aljaafari, Nura, et al.
Publicado: (2026)
Estimating the Causal Effects of Natural Logic Features in Transformer-Based NLI Models
por: Rozanova, Julia, et al.
Publicado: (2024)
por: Rozanova, Julia, et al.
Publicado: (2024)
SylloBio-NLI: Evaluating Large Language Models on Biomedical Syllogistic Reasoning
por: Wysocka, Magdalena, et al.
Publicado: (2024)
por: Wysocka, Magdalena, et al.
Publicado: (2024)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
por: Quan, Xin, et al.
Publicado: (2025)
por: Quan, Xin, et al.
Publicado: (2025)
SemEval-2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials
por: Jullien, Mael, et al.
Publicado: (2024)
por: Jullien, Mael, et al.
Publicado: (2024)
Multi-Relational Hyperbolic Word Embeddings from Natural Language Definitions
por: Valentino, Marco, et al.
Publicado: (2023)
por: Valentino, Marco, et al.
Publicado: (2023)
FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean
por: Meadows, Jordan, et al.
Publicado: (2026)
por: Meadows, Jordan, et al.
Publicado: (2026)
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
por: Valentino, Marco, et al.
Publicado: (2025)
por: Valentino, Marco, et al.
Publicado: (2025)
Multi-Operational Mathematical Derivations in Latent Space
por: Valentino, Marco, et al.
Publicado: (2023)
por: Valentino, Marco, et al.
Publicado: (2023)
A Differentiable Integer Linear Programming Solver for Explanation-Based Natural Language Inference
por: Thayaparan, Mokanarangan, et al.
Publicado: (2024)
por: Thayaparan, Mokanarangan, et al.
Publicado: (2024)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
por: Lee, Dongryeol, et al.
Publicado: (2024)
por: Lee, Dongryeol, et al.
Publicado: (2024)
Decompose-and-Formalise: Recursively Verifiable Natural Language Inference
por: Quan, Xin, et al.
Publicado: (2026)
por: Quan, Xin, et al.
Publicado: (2026)
Enhancing Ethical Explanations of Large Language Models through Iterative Symbolic Refinement
por: Quan, Xin, et al.
Publicado: (2024)
por: Quan, Xin, et al.
Publicado: (2024)
Improving Semantic Control in Discrete Latent Spaces with Transformer Quantized Variational Autoencoders
por: Zhang, Yingji, et al.
Publicado: (2024)
por: Zhang, Yingji, et al.
Publicado: (2024)
Conjecturing: An Overlooked Step in Formal Mathematical Reasoning
por: Sivakumar, Jasivan Alex, et al.
Publicado: (2025)
por: Sivakumar, Jasivan Alex, et al.
Publicado: (2025)
A Survey in Mathematical Language Processing
por: Meadows, Jordan, et al.
Publicado: (2022)
por: Meadows, Jordan, et al.
Publicado: (2022)
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
por: Zhang, Jiaqiao, et al.
Publicado: (2026)
por: Zhang, Jiaqiao, et al.
Publicado: (2026)
Evaluating Mathematical Reasoning Beyond Accuracy
por: Xia, Shijie, et al.
Publicado: (2024)
por: Xia, Shijie, et al.
Publicado: (2024)
Integrating Expert Knowledge into Logical Programs via LLMs
por: Górski, Franciszek, et al.
Publicado: (2025)
por: Górski, Franciszek, et al.
Publicado: (2025)
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability
por: Wang, Ruida, et al.
Publicado: (2025)
por: Wang, Ruida, et al.
Publicado: (2025)
Inference to the Best Explanation in Large Language Models
por: Dalal, Dhairya, et al.
Publicado: (2024)
por: Dalal, Dhairya, et al.
Publicado: (2024)
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
por: Bourigault, Pauline, et al.
Publicado: (2026)
por: Bourigault, Pauline, et al.
Publicado: (2026)
Ejemplares similares
-
Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions
por: Zhang, Lan, et al.
Publicado: (2025) -
MASA: LLM-Driven Multi-Agent Systems for Autoformalization
por: Zhang, Lan, et al.
Publicado: (2025) -
PEIRCE: Unifying Material and Formal Reasoning via LLM-Driven Neuro-Symbolic Refinement
por: Quan, Xin, et al.
Publicado: (2025) -
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
por: Meadows, Jordan, et al.
Publicado: (2023) -
Reasoning with Natural Language Explanations
por: Valentino, Marco, et al.
Publicado: (2024)