Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xiao, Wu, Zirui, Wu, Xueqing, Lu, Pan, Chang, Kai-Wei, Feng, Yansong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
von: Liu, Xiao, et al.
Veröffentlicht: (2025)
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
von: Lin, Jiuheng, et al.
Veröffentlicht: (2025)
von: Lin, Jiuheng, et al.
Veröffentlicht: (2025)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
Explore the Reasoning Capability of LLMs in the Chess Testbed
von: Wang, Shu, et al.
Veröffentlicht: (2024)
von: Wang, Shu, et al.
Veröffentlicht: (2024)
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
von: Lai, Siqi, et al.
Veröffentlicht: (2025)
von: Lai, Siqi, et al.
Veröffentlicht: (2025)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
von: Deng, Wenhao, et al.
Veröffentlicht: (2025)
von: Deng, Wenhao, et al.
Veröffentlicht: (2025)
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
von: Alshaikh, Rana, et al.
Veröffentlicht: (2025)
von: Alshaikh, Rana, et al.
Veröffentlicht: (2025)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
von: Hong, Zijin, et al.
Veröffentlicht: (2025)
von: Hong, Zijin, et al.
Veröffentlicht: (2025)
CausalEval: Towards Better Causal Reasoning in Language Models
von: Yu, Longxuan, et al.
Veröffentlicht: (2024)
von: Yu, Longxuan, et al.
Veröffentlicht: (2024)
Reasoning about Affordances: Causal and Compositional Reasoning in LLMs
von: Gjerde, Magnus F., et al.
Veröffentlicht: (2025)
von: Gjerde, Magnus F., et al.
Veröffentlicht: (2025)
CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention
von: Sun, Yuxi, et al.
Veröffentlicht: (2025)
von: Sun, Yuxi, et al.
Veröffentlicht: (2025)
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
von: Kumarage, Tharindu, et al.
Veröffentlicht: (2025)
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
von: Feng, Guhao, et al.
Veröffentlicht: (2024)
von: Feng, Guhao, et al.
Veröffentlicht: (2024)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models
von: Komanduri, Aneesh, et al.
Veröffentlicht: (2025)
von: Komanduri, Aneesh, et al.
Veröffentlicht: (2025)
How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning
von: Chen, Haoyang, et al.
Veröffentlicht: (2026)
von: Chen, Haoyang, et al.
Veröffentlicht: (2026)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
von: Chen, Simin, et al.
Veröffentlicht: (2025)
von: Chen, Simin, et al.
Veröffentlicht: (2025)
CreDes: Causal Reasoning Enhancement and Dual-End Searching for Solving Long-Range Reasoning Problems using LLMs
von: Wang, Kangsheng, et al.
Veröffentlicht: (2024)
von: Wang, Kangsheng, et al.
Veröffentlicht: (2024)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
ProTrix: Building Models for Planning and Reasoning over Tables with Sentence Context
von: Wu, Zirui, et al.
Veröffentlicht: (2024)
von: Wu, Zirui, et al.
Veröffentlicht: (2024)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
von: Lei, Chao, et al.
Veröffentlicht: (2025)
von: Lei, Chao, et al.
Veröffentlicht: (2025)
Kinship Data Benchmark for Multi-hop Reasoning
von: Sun, Tianda, et al.
Veröffentlicht: (2026)
von: Sun, Tianda, et al.
Veröffentlicht: (2026)
Com$^2$: A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models
von: Xiong, Kai, et al.
Veröffentlicht: (2025)
von: Xiong, Kai, et al.
Veröffentlicht: (2025)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
von: Wu, Junhong, et al.
Veröffentlicht: (2025)
von: Wu, Junhong, et al.
Veröffentlicht: (2025)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
The Art of Efficient Reasoning: Data, Reward, and Optimization
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2026)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
von: Shi, Dachuan, et al.
Veröffentlicht: (2025)
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
von: Yang, Chang, et al.
Veröffentlicht: (2025)
von: Yang, Chang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning
von: Liu, Xiao, et al.
Veröffentlicht: (2025) -
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
von: Lin, Jiuheng, et al.
Veröffentlicht: (2025) -
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024) -
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
von: Wu, Xueqing, et al.
Veröffentlicht: (2024) -
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)