BaRDa: A Belief and Reasoning Dataset that Separates Factual Accuracy and Reasoning Ability
Fuente:
arXiv
Saved in:
| Main Authors: | Clark, Peter, Mishra, Bhavana Dalvi, Tafjord, Oyvind |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
by: Jansen, Peter, et al.
Published: (2024)
by: Jansen, Peter, et al.
Published: (2024)
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
by: Jansen, Peter, et al.
Published: (2025)
by: Jansen, Peter, et al.
Published: (2025)
Digital Socrates: Evaluating LLMs through Explanation Critiques
by: Gu, Yuling, et al.
Published: (2023)
by: Gu, Yuling, et al.
Published: (2023)
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
by: Weir, Nathaniel, et al.
Published: (2024)
by: Weir, Nathaniel, et al.
Published: (2024)
HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance
by: Vasu, Rosni, et al.
Published: (2025)
by: Vasu, Rosni, et al.
Published: (2025)
HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation
by: Vasu, Rosni, et al.
Published: (2025)
by: Vasu, Rosni, et al.
Published: (2025)
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
by: Gu, Yuling, et al.
Published: (2024)
by: Gu, Yuling, et al.
Published: (2024)
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
by: Jiao, Rui, et al.
Published: (2025)
by: Jiao, Rui, et al.
Published: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question Answering
by: Weir, Nathaniel, et al.
Published: (2024)
by: Weir, Nathaniel, et al.
Published: (2024)
OLMES: A Standard for Language Model Evaluations
by: Gu, Yuling, et al.
Published: (2024)
by: Gu, Yuling, et al.
Published: (2024)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
by: Wu, Juncheng, et al.
Published: (2025)
by: Wu, Juncheng, et al.
Published: (2025)
Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning Models
by: Li, Junyi, et al.
Published: (2025)
by: Li, Junyi, et al.
Published: (2025)
Reasoning Factual Knowledge in Structured Data with Large Language Models
by: Huang, Sirui, et al.
Published: (2024)
by: Huang, Sirui, et al.
Published: (2024)
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
by: Mishra, Prakamya, et al.
Published: (2025)
by: Mishra, Prakamya, et al.
Published: (2025)
Diversity of Thought Improves Reasoning Abilities of LLMs
by: Naik, Ranjita, et al.
Published: (2023)
by: Naik, Ranjita, et al.
Published: (2023)
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
by: Majumder, Bodhisattwa Prasad, et al.
Published: (2024)
ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation
by: Lee, Zhicheng, et al.
Published: (2025)
by: Lee, Zhicheng, et al.
Published: (2025)
Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering
by: Leong, Chak Tou, et al.
Published: (2026)
by: Leong, Chak Tou, et al.
Published: (2026)
Aligning Knowledge Graphs and Language Models for Factual Accuracy
by: Nishat, Nur A Zarin, et al.
Published: (2025)
by: Nishat, Nur A Zarin, et al.
Published: (2025)
Editing Factual Knowledge and Explanatory Ability of Medical Large Language Models
by: Xu, Derong, et al.
Published: (2024)
by: Xu, Derong, et al.
Published: (2024)
Reasoning Does Not Necessarily Improve Role-Playing Ability
by: Feng, Xiachong, et al.
Published: (2025)
by: Feng, Xiachong, et al.
Published: (2025)
Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
by: Cai, Wenrui, et al.
Published: (2025)
by: Cai, Wenrui, et al.
Published: (2025)
BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs
by: Yang, Junxiao, et al.
Published: (2025)
by: Yang, Junxiao, et al.
Published: (2025)
K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning
by: Zhang, Yadong, et al.
Published: (2024)
by: Zhang, Yadong, et al.
Published: (2024)
FEABench: Evaluating Language Models on Multiphysics Reasoning Ability
by: Mudur, Nayantara, et al.
Published: (2025)
by: Mudur, Nayantara, et al.
Published: (2025)
A Survey on Enhancing Causal Reasoning Ability of Large Language Models
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
FactReasoner: A Probabilistic Approach to Long-Form Factuality Assessment for Large Language Models
by: Marinescu, Radu, et al.
Published: (2025)
by: Marinescu, Radu, et al.
Published: (2025)
MDCR: A Dataset for Multi-Document Conditional Reasoning
by: Chen, Peter Baile, et al.
Published: (2024)
by: Chen, Peter Baile, et al.
Published: (2024)
BIS Reasoning 1.0: The First Large-Scale Japanese Benchmark for Belief-Inconsistent Syllogistic Reasoning
by: Nguyen, Ha-Thanh, et al.
Published: (2025)
by: Nguyen, Ha-Thanh, et al.
Published: (2025)
Eliciting Causal Abilities in Large Language Models for Reasoning Tasks
by: Wang, Yajing, et al.
Published: (2024)
by: Wang, Yajing, et al.
Published: (2024)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
by: Xu, Jillian, et al.
Published: (2025)
by: Xu, Jillian, et al.
Published: (2025)
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
by: Wang, Yiqi, et al.
Published: (2024)
by: Wang, Yiqi, et al.
Published: (2024)
AdaMCoT: Rethinking Cross-Lingual Factual Reasoning through Adaptive Multilingual Chain-of-Thought
by: Zheng, Weihua, et al.
Published: (2025)
by: Zheng, Weihua, et al.
Published: (2025)
Follow the Path: Reasoning over Knowledge Graph Paths to Improve Large Language Model Factuality
by: Zhang, Mike, et al.
Published: (2025)
by: Zhang, Mike, et al.
Published: (2025)
Learning to Reason via Program Generation, Emulation, and Search
by: Weir, Nathaniel, et al.
Published: (2024)
by: Weir, Nathaniel, et al.
Published: (2024)
Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Similar Items
-
DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
by: Jansen, Peter, et al.
Published: (2024) -
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
by: Jansen, Peter, et al.
Published: (2025) -
Digital Socrates: Evaluating LLMs through Explanation Critiques
by: Gu, Yuling, et al.
Published: (2023) -
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
by: Weir, Nathaniel, et al.
Published: (2024) -
HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance
by: Vasu, Rosni, et al.
Published: (2025)