DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Pandey, Atharva, Dubey, Kshitij, Sharma, Rahul, Sharma, Amit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
di: Xu, Xinnuo, et al.
Pubblicazione: (2025)
di: Xu, Xinnuo, et al.
Pubblicazione: (2025)
Teaching Transformers Causal Reasoning through Axiomatic Training
di: Vashishtha, Aniket, et al.
Pubblicazione: (2024)
di: Vashishtha, Aniket, et al.
Pubblicazione: (2024)
The Role of Deductive and Inductive Reasoning in Large Language Models
di: Cai, Chengkun, et al.
Pubblicazione: (2024)
di: Cai, Chengkun, et al.
Pubblicazione: (2024)
NICE: To Optimize In-Context Examples or Not?
di: Srivastava, Pragya, et al.
Pubblicazione: (2024)
di: Srivastava, Pragya, et al.
Pubblicazione: (2024)
Temporal Consistency for LLM Reasoning Process Error Identification
di: Guo, Jiacheng, et al.
Pubblicazione: (2025)
di: Guo, Jiacheng, et al.
Pubblicazione: (2025)
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
di: Chen, Michael K., et al.
Pubblicazione: (2025)
di: Chen, Michael K., et al.
Pubblicazione: (2025)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
di: Lee, Jaehyeok, et al.
Pubblicazione: (2024)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2024)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
di: Zhou, Zhi, et al.
Pubblicazione: (2025)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
di: Wu, Da, et al.
Pubblicazione: (2023)
di: Wu, Da, et al.
Pubblicazione: (2023)
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
di: Xu, Yifei, et al.
Pubblicazione: (2025)
di: Xu, Yifei, et al.
Pubblicazione: (2025)
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
di: Kıcıman, Emre, et al.
Pubblicazione: (2023)
di: Kıcıman, Emre, et al.
Pubblicazione: (2023)
SciNets: Graph-Constrained Multi-Hop Reasoning for Scientific Literature Synthesis
di: Dubey, Sauhard
Pubblicazione: (2025)
di: Dubey, Sauhard
Pubblicazione: (2025)
ARGUS: Adaptive Rotation-Invariant Geometric Unsupervised System
di: Sharma, Anantha
Pubblicazione: (2026)
di: Sharma, Anantha
Pubblicazione: (2026)
Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
di: Yousuf, Raquib Bin, et al.
Pubblicazione: (2025)
di: Yousuf, Raquib Bin, et al.
Pubblicazione: (2025)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
di: Nath, Vaskar, et al.
Pubblicazione: (2025)
di: Nath, Vaskar, et al.
Pubblicazione: (2025)
GEAR: A General Evaluation Framework for Abductive Reasoning
di: He, Kaiyu, et al.
Pubblicazione: (2025)
di: He, Kaiyu, et al.
Pubblicazione: (2025)
Task Facet Learning: A Structured Approach to Prompt Optimization
di: Juneja, Gurusha, et al.
Pubblicazione: (2024)
di: Juneja, Gurusha, et al.
Pubblicazione: (2024)
Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation
di: Bao, Qiming, et al.
Pubblicazione: (2022)
di: Bao, Qiming, et al.
Pubblicazione: (2022)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
di: Kim, Eunsu, et al.
Pubblicazione: (2025)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
Towards Optimizing the Costs of LLM Usage
di: Shekhar, Shivanshu, et al.
Pubblicazione: (2024)
di: Shekhar, Shivanshu, et al.
Pubblicazione: (2024)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
di: Kulkarni, Atharva, et al.
Pubblicazione: (2025)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2025)
ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
di: Banyas, Peter, et al.
Pubblicazione: (2025)
di: Banyas, Peter, et al.
Pubblicazione: (2025)
CoLa: Learning to Interactively Collaborate with Large Language Models
di: Sharma, Abhishek, et al.
Pubblicazione: (2025)
di: Sharma, Abhishek, et al.
Pubblicazione: (2025)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics
di: Sharma, Kartik, et al.
Pubblicazione: (2026)
di: Sharma, Kartik, et al.
Pubblicazione: (2026)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
di: Potamitis, Nearchos, et al.
Pubblicazione: (2025)
di: Potamitis, Nearchos, et al.
Pubblicazione: (2025)
Stepwise Self-Consistent Mathematical Reasoning with Large Language Models
di: Zhao, Zilong, et al.
Pubblicazione: (2024)
di: Zhao, Zilong, et al.
Pubblicazione: (2024)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
di: Muhamed, Aashiq
Pubblicazione: (2025)
di: Muhamed, Aashiq
Pubblicazione: (2025)
LIME-LLM: Probing Models with Fluent Counterfactuals, Not Broken Text
di: Mihaila, George, et al.
Pubblicazione: (2026)
di: Mihaila, George, et al.
Pubblicazione: (2026)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
di: Sharma, Archit, et al.
Pubblicazione: (2024)
di: Sharma, Archit, et al.
Pubblicazione: (2024)
Finding the Cracks: Improving LLMs Reasoning with Paraphrastic Probing and Consistency Verification
di: Shi, Weili, et al.
Pubblicazione: (2026)
di: Shi, Weili, et al.
Pubblicazione: (2026)
Context-Enhanced Contrastive Search for Improved LLM Text Generation
di: Sen, Jaydip, et al.
Pubblicazione: (2025)
di: Sen, Jaydip, et al.
Pubblicazione: (2025)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
di: Sansford, Hannah, et al.
Pubblicazione: (2024)
di: Sansford, Hannah, et al.
Pubblicazione: (2024)
LegalSeg: Unlocking the Structure of Indian Legal Judgments Through Rhetorical Role Classification
di: Nigam, Shubham Kumar, et al.
Pubblicazione: (2025)
di: Nigam, Shubham Kumar, et al.
Pubblicazione: (2025)
Less is More for Improving Automatic Evaluation of Factual Consistency
di: Wang, Tong, et al.
Pubblicazione: (2024)
di: Wang, Tong, et al.
Pubblicazione: (2024)
SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection
di: Hu, Mengya, et al.
Pubblicazione: (2024)
di: Hu, Mengya, et al.
Pubblicazione: (2024)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
di: Yoa, Seungdong, et al.
Pubblicazione: (2026)
di: Yoa, Seungdong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
di: Xu, Xinnuo, et al.
Pubblicazione: (2025) -
Teaching Transformers Causal Reasoning through Axiomatic Training
di: Vashishtha, Aniket, et al.
Pubblicazione: (2024) -
The Role of Deductive and Inductive Reasoning in Large Language Models
di: Cai, Chengkun, et al.
Pubblicazione: (2024) -
NICE: To Optimize In-Context Examples or Not?
di: Srivastava, Pragya, et al.
Pubblicazione: (2024) -
Temporal Consistency for LLM Reasoning Process Error Identification
di: Guo, Jiacheng, et al.
Pubblicazione: (2025)