Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aggarwal, Shashank, Mishra, Ram Vikas, Awekar, Amit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Reasoning Chains through Depth-Specialized Mixture-of-Experts in Transformer Architectures
von: Roy, Sampurna, et al.
Veröffentlicht: (2025)
von: Roy, Sampurna, et al.
Veröffentlicht: (2025)
GRACE: Generative Recommendation via Journey-Aware Sparse Attention on Chain-of-Thought Tokenization
von: Ma, Luyi, et al.
Veröffentlicht: (2025)
von: Ma, Luyi, et al.
Veröffentlicht: (2025)
Are Word Embedding Methods Stable and Should We Care About It?
von: Borah, Angana, et al.
Veröffentlicht: (2021)
von: Borah, Angana, et al.
Veröffentlicht: (2021)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
von: Wang, Guangzhi, et al.
Veröffentlicht: (2025)
von: Wang, Guangzhi, et al.
Veröffentlicht: (2025)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
von: Pradeep, Ronak, et al.
Veröffentlicht: (2024)
von: Pradeep, Ronak, et al.
Veröffentlicht: (2024)
Graph-DPEP: Decomposed Plug and Ensemble Play for Few-Shot Document Relation Extraction with Graph-of-Thoughts Reasoning
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
Automated Extraction and Creation of FBS Design Reasoning Knowledge Graphs from Structured Data in Product Catalogues Lacking Contextual Information
von: Sahadevan, Vijayalaxmi, et al.
Veröffentlicht: (2024)
von: Sahadevan, Vijayalaxmi, et al.
Veröffentlicht: (2024)
Alignment Adapter to Improve the Performance of Compressed Deep Learning Models
von: Rai, Rohit Raj, et al.
Veröffentlicht: (2026)
von: Rai, Rohit Raj, et al.
Veröffentlicht: (2026)
An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs
von: Ma, Linyue, et al.
Veröffentlicht: (2025)
von: Ma, Linyue, et al.
Veröffentlicht: (2025)
Thinking Broad, Acting Fast: Latent Reasoning Distillation from Multi-Perspective Chain-of-Thought for E-Commerce Relevance
von: Qiu, Baopu, et al.
Veröffentlicht: (2026)
von: Qiu, Baopu, et al.
Veröffentlicht: (2026)
DoTA-RAG: Dynamic of Thought Aggregation RAG
von: Ruangtanusak, Saksorn, et al.
Veröffentlicht: (2025)
von: Ruangtanusak, Saksorn, et al.
Veröffentlicht: (2025)
Thought-Augmented Planning for LLM-Powered Interactive Recommender Agent
von: Yu, Haocheng, et al.
Veröffentlicht: (2025)
von: Yu, Haocheng, et al.
Veröffentlicht: (2025)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
von: Thakur, Nandan, et al.
Veröffentlicht: (2026)
von: Thakur, Nandan, et al.
Veröffentlicht: (2026)
OpenNovelty: An LLM-powered Agentic System for Verifiable Scholarly Novelty Assessment
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
Reason to Contrast: A Cascaded Multimodal Retrieval Framework
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
von: Cui, Xuanming, et al.
Veröffentlicht: (2025)
VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering
von: Košprdić, Miloš, et al.
Veröffentlicht: (2026)
von: Košprdić, Miloš, et al.
Veröffentlicht: (2026)
OrgForge: A Multi-Agent Simulation Framework for Verifiable Synthetic Corporate Corpora
von: Flynt, Jeffrey
Veröffentlicht: (2026)
von: Flynt, Jeffrey
Veröffentlicht: (2026)
Improving Medical Reasoning through Retrieval and Self-Reflection with Retrieval-Augmented Large Language Models
von: Jeong, Minbyul, et al.
Veröffentlicht: (2024)
von: Jeong, Minbyul, et al.
Veröffentlicht: (2024)
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
von: Tan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Tan, Zhiwen, et al.
Veröffentlicht: (2025)
Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
von: Wu, Junlin, et al.
Veröffentlicht: (2025)
von: Wu, Junlin, et al.
Veröffentlicht: (2025)
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
von: Huang, Youcheng, et al.
Veröffentlicht: (2025)
TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question Answering
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
von: Markowitz, Elan, et al.
Veröffentlicht: (2025)
Tutorial on Reasoning for IR & IR for Reasoning
von: Hoveyda, Mohanna, et al.
Veröffentlicht: (2026)
von: Hoveyda, Mohanna, et al.
Veröffentlicht: (2026)
Text-to-SPARQL Goes Beyond English: Multilingual Question Answering Over Knowledge Graphs through Human-Inspired Reasoning
von: Perevalov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Perevalov, Aleksandr, et al.
Veröffentlicht: (2025)
Chained Prompting for Better Systematic Review Search Strategies
von: Nasser, Fatima, et al.
Veröffentlicht: (2025)
von: Nasser, Fatima, et al.
Veröffentlicht: (2025)
From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
von: Mishra, Shubham, et al.
Veröffentlicht: (2025)
von: Mishra, Shubham, et al.
Veröffentlicht: (2025)
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
von: Islamaj, Rezarta, et al.
Veröffentlicht: (2026)
von: Islamaj, Rezarta, et al.
Veröffentlicht: (2026)
RAG-KG-IL: A Multi-Agent Hybrid Framework for Reducing Hallucinations and Enhancing LLM Reasoning through RAG and Incremental Knowledge Graph Learning Integration
von: Yu, Hong Qing, et al.
Veröffentlicht: (2025)
von: Yu, Hong Qing, et al.
Veröffentlicht: (2025)
Enhancing Supply Chain Visibility with Knowledge Graphs and Large Language Models
von: AlMahri, Sara, et al.
Veröffentlicht: (2024)
von: AlMahri, Sara, et al.
Veröffentlicht: (2024)
LimGen: Probing the LLMs for Generating Suggestive Limitations of Research Papers
von: Faizullah, Abdur Rahman Bin Md, et al.
Veröffentlicht: (2024)
von: Faizullah, Abdur Rahman Bin Md, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Reasoning for Chartered Accountancy
von: Gupta, Jatin, et al.
Veröffentlicht: (2026)
von: Gupta, Jatin, et al.
Veröffentlicht: (2026)
On Reasoning Behind Next Occupation Recommendation
von: Dong, Shan, et al.
Veröffentlicht: (2026)
von: Dong, Shan, et al.
Veröffentlicht: (2026)
Verify-in-the-Graph: Entity Disambiguation Enhancement for Complex Claim Verification with Interactive Graph Representation
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
Fast and Accurate Contextual Knowledge Extraction Using Cascading Language Model Chains and Candidate Answers
von: Harris, Lee
Veröffentlicht: (2025)
von: Harris, Lee
Veröffentlicht: (2025)
ALARB: An Arabic Legal Argument Reasoning Benchmark
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
von: Shairah, Harethah Abu, et al.
Veröffentlicht: (2025)
Combining Evidence and Reasoning for Biomedical Fact-Checking
von: Barone, Mariano, et al.
Veröffentlicht: (2025)
von: Barone, Mariano, et al.
Veröffentlicht: (2025)
HORIZON: A Benchmark for In-the-wild User Behaviour Modeling
von: Goel, Arnav, et al.
Veröffentlicht: (2026)
von: Goel, Arnav, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dynamic Reasoning Chains through Depth-Specialized Mixture-of-Experts in Transformer Architectures
von: Roy, Sampurna, et al.
Veröffentlicht: (2025) -
GRACE: Generative Recommendation via Journey-Aware Sparse Attention on Chain-of-Thought Tokenization
von: Ma, Luyi, et al.
Veröffentlicht: (2025) -
Are Word Embedding Methods Stable and Should We Care About It?
von: Borah, Angana, et al.
Veröffentlicht: (2021) -
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
von: Li, Xiaonan, et al.
Veröffentlicht: (2023) -
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
von: Wang, Guangzhi, et al.
Veröffentlicht: (2025)