Evaluating Self: Consistency and Tree of Thought Reasoning Strategies in Lightweight Language Models on Math and Logic Benchmark
Fuente:
Zenodo
Gespeichert in:
| Hauptverfasser: | Mr. Sankhadeep Debdas, Pawan Kumar |
|---|---|
| Format: | Recurso digital |
| Veröffentlicht: |
Zenodo
2026
|
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems
von: Kai, Ding, et al.
Veröffentlicht: (2024)
von: Kai, Ding, et al.
Veröffentlicht: (2024)
Bridging the Black Box: An Interpretable Deep Learning Framework for Multi-Class Dermatological Screening
von: Mr. Aniket Kumar, et al.
Veröffentlicht: (2026)
von: Mr. Aniket Kumar, et al.
Veröffentlicht: (2026)
Self-Consistency Boosts Calibration for Math Reasoning
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning
von: Toh, Vernon Y. H., et al.
Veröffentlicht: (2024)
von: Toh, Vernon Y. H., et al.
Veröffentlicht: (2024)
ProofSketcher: Hybrid LLM + Lightweight Proof Checker for Reliable Math/Logic Reasoning
von: Kommuru, Kranthi, et al.
Veröffentlicht: (2026)
von: Kommuru, Kranthi, et al.
Veröffentlicht: (2026)
Benchmarking Large Language Models for Math Reasoning Tasks
von: Seßler, Kathrin, et al.
Veröffentlicht: (2024)
von: Seßler, Kathrin, et al.
Veröffentlicht: (2024)
Structured Reasoning with Tree-of-Thoughts for Bengali Math Word Problems
von: Mahmood, Aurprita, et al.
Veröffentlicht: (2025)
von: Mahmood, Aurprita, et al.
Veröffentlicht: (2025)
Scheherazade: Evaluating Chain-of-Thought Math Reasoning in LLMs with Chain-of-Problems
von: Miner, Stephen, et al.
Veröffentlicht: (2024)
von: Miner, Stephen, et al.
Veröffentlicht: (2024)
UTMath: Math Evaluation with Unit Test via Reasoning-to-Coding Thoughts
von: Yang, Bo, et al.
Veröffentlicht: (2024)
von: Yang, Bo, et al.
Veröffentlicht: (2024)
Logic-of-Thought: Injecting Logic into Contexts for Full Reasoning in Large Language Models
von: Liu, Tongxuan, et al.
Veröffentlicht: (2024)
von: Liu, Tongxuan, et al.
Veröffentlicht: (2024)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
von: Li, Chengpeng, et al.
Veröffentlicht: (2024)
von: Li, Chengpeng, et al.
Veröffentlicht: (2024)
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
von: Feng, Jun, et al.
Veröffentlicht: (2025)
von: Feng, Jun, et al.
Veröffentlicht: (2025)
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
von: Liu, Junming, et al.
Veröffentlicht: (2026)
von: Liu, Junming, et al.
Veröffentlicht: (2026)
Inference-Time Rethinking with Latent Thought Vectors for Math Reasoning
von: Kong, Deqian, et al.
Veröffentlicht: (2026)
von: Kong, Deqian, et al.
Veröffentlicht: (2026)
Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning
von: Paul, Bidyarthi, et al.
Veröffentlicht: (2025)
von: Paul, Bidyarthi, et al.
Veröffentlicht: (2025)
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2025)
Cubic Regularization Technique of the Newton Method for Vector Optimization
von: Ghosh, Debdas
Veröffentlicht: (2025)
von: Ghosh, Debdas
Veröffentlicht: (2025)
VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2023)
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2023)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
Towards Logically Consistent Language Models via Probabilistic Reasoning
von: Calanzone, Diego, et al.
Veröffentlicht: (2024)
von: Calanzone, Diego, et al.
Veröffentlicht: (2024)
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
von: Sun, Haoxiang, et al.
Veröffentlicht: (2025)
von: Sun, Haoxiang, et al.
Veröffentlicht: (2025)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2025)
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2025)
LogicCat: A Chain-of-Thought Text-to-SQL Benchmark for Complex Reasoning
von: Liu, Tao, et al.
Veröffentlicht: (2025)
von: Liu, Tao, et al.
Veröffentlicht: (2025)
AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
von: Wen, Yibin, et al.
Veröffentlicht: (2025)
von: Wen, Yibin, et al.
Veröffentlicht: (2025)
SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models
von: Xie, Yuechen, et al.
Veröffentlicht: (2026)
von: Xie, Yuechen, et al.
Veröffentlicht: (2026)
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
von: Glazer, Elliot, et al.
Veröffentlicht: (2024)
von: Glazer, Elliot, et al.
Veröffentlicht: (2024)
DCR-Consistency: Divide-Conquer-Reasoning for Consistency Evaluation and Improvement of Large Language Models
von: Cui, Wendi, et al.
Veröffentlicht: (2024)
von: Cui, Wendi, et al.
Veröffentlicht: (2024)
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2025)
Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2025)
von: Mao, Zhenjiang, et al.
Veröffentlicht: (2025)
Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models
von: Liu, Yinhong, et al.
Veröffentlicht: (2024)
von: Liu, Yinhong, et al.
Veröffentlicht: (2024)
Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic
von: Zhao, Xufeng, et al.
Veröffentlicht: (2023)
von: Zhao, Xufeng, et al.
Veröffentlicht: (2023)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2025)
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2025)
STEM-POM: Evaluating Language Models Math-Symbol Reasoning in Document Parsing
von: Zou, Jiaru, et al.
Veröffentlicht: (2024)
von: Zou, Jiaru, et al.
Veröffentlicht: (2024)
Mathfish: Evaluating Language Model Math Reasoning via Grounding in Educational Curricula
von: Lucy, Li, et al.
Veröffentlicht: (2024)
von: Lucy, Li, et al.
Veröffentlicht: (2024)
Learning to Reason via Mixture-of-Thought for Logical Reasoning
von: Zheng, Tong, et al.
Veröffentlicht: (2025)
von: Zheng, Tong, et al.
Veröffentlicht: (2025)
ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
von: Tanmay, Kumar, et al.
Veröffentlicht: (2025)
von: Tanmay, Kumar, et al.
Veröffentlicht: (2025)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems
von: Kai, Ding, et al.
Veröffentlicht: (2024) -
Bridging the Black Box: An Interpretable Deep Learning Framework for Multi-Class Dermatological Screening
von: Mr. Aniket Kumar, et al.
Veröffentlicht: (2026) -
Self-Consistency Boosts Calibration for Math Reasoning
von: Wang, Ante, et al.
Veröffentlicht: (2024) -
Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning
von: Toh, Vernon Y. H., et al.
Veröffentlicht: (2024) -
ProofSketcher: Hybrid LLM + Lightweight Proof Checker for Reliable Math/Logic Reasoning
von: Kommuru, Kranthi, et al.
Veröffentlicht: (2026)