Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Javaji, Shashidhar Reddy, Cao, Yupeng, Li, Haohang, Yu, Yangyang, Muralidhar, Nikhil, Zhu, Zining |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2024)
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2024)
LLM-Generated Black-box Explanations Can Be Adversarially Helpful
von: Ajwani, Rohan, et al.
Veröffentlicht: (2024)
von: Ajwani, Rohan, et al.
Veröffentlicht: (2024)
Truth Neurons
von: Li, Haohang, et al.
Veröffentlicht: (2025)
von: Li, Haohang, et al.
Veröffentlicht: (2025)
VERBA: Verbalizing Model Differences Using Large Language Models
von: Doda, Shravan, et al.
Veröffentlicht: (2025)
von: Doda, Shravan, et al.
Veröffentlicht: (2025)
Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2025)
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2025)
INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent
von: Li, Haohang, et al.
Veröffentlicht: (2024)
von: Li, Haohang, et al.
Veröffentlicht: (2024)
CiteME: Can Language Models Accurately Cite Scientific Claims?
von: Press, Ori, et al.
Veröffentlicht: (2024)
von: Press, Ori, et al.
Veröffentlicht: (2024)
FinAudio: A Benchmark for Audio Large Language Models in Financial Applications
von: Cao, Yupeng, et al.
Veröffentlicht: (2025)
von: Cao, Yupeng, et al.
Veröffentlicht: (2025)
Assessing the Reasoning Capabilities of LLMs in the context of Evidence-based Claim Verification
von: Dougrez-Lewis, John, et al.
Veröffentlicht: (2024)
von: Dougrez-Lewis, John, et al.
Veröffentlicht: (2024)
Atomic Reasoning for Scientific Table Claim Verification
von: Zhang, Yuji, et al.
Veröffentlicht: (2025)
von: Zhang, Yuji, et al.
Veröffentlicht: (2025)
Automated ensemble method for pediatric brain tumor segmentation
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2023)
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2023)
SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification
von: Kumar, Sujit, et al.
Veröffentlicht: (2025)
von: Kumar, Sujit, et al.
Veröffentlicht: (2025)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
von: James, Joseph, et al.
Veröffentlicht: (2026)
von: James, Joseph, et al.
Veröffentlicht: (2026)
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
von: Ge, Huaizhi, et al.
Veröffentlicht: (2024)
MuSciClaims: Multimodal Scientific Claim Verification
von: Lal, Yash Kumar, et al.
Veröffentlicht: (2025)
von: Lal, Yash Kumar, et al.
Veröffentlicht: (2025)
ClaimFlow: Tracing the Evolution of Scientific Claims in NLP
von: Pramanick, Aniket, et al.
Veröffentlicht: (2026)
von: Pramanick, Aniket, et al.
Veröffentlicht: (2026)
FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks
von: Cao, Yupeng, et al.
Veröffentlicht: (2026)
von: Cao, Yupeng, et al.
Veröffentlicht: (2026)
MERMAID: Memory-Enhanced Retrieval and Reasoning with Multi-Agent Iterative Knowledge Grounding for Veracity Assessment
von: Cao, Yupeng, et al.
Veröffentlicht: (2026)
von: Cao, Yupeng, et al.
Veröffentlicht: (2026)
Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce
von: Wang, Haojin, et al.
Veröffentlicht: (2025)
von: Wang, Haojin, et al.
Veröffentlicht: (2025)
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
von: Ho, Xanh, et al.
Veröffentlicht: (2026)
von: Ho, Xanh, et al.
Veröffentlicht: (2026)
Can Large Language Models Detect Misinformation in Scientific News Reporting?
von: Cao, Yupeng, et al.
Veröffentlicht: (2024)
von: Cao, Yupeng, et al.
Veröffentlicht: (2024)
On Higher Order Busy Beaver Function
von: Cao, Zining
Veröffentlicht: (2025)
von: Cao, Zining
Veröffentlicht: (2025)
Deceptive Humor: A Synthetic Multilingual Benchmark Dataset for Bridging Fabricated Claims with Humorous Content
von: Kasu, Sai Kartheek Reddy, et al.
Veröffentlicht: (2025)
von: Kasu, Sai Kartheek Reddy, et al.
Veröffentlicht: (2025)
Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences
von: Koneru, Sai, et al.
Veröffentlicht: (2023)
von: Koneru, Sai, et al.
Veröffentlicht: (2023)
Do Gender Cues Affect LLM Value Trade-offs? Evidence from a Controlled Decision Benchmark
von: Liu, Yangyang, et al.
Veröffentlicht: (2026)
von: Liu, Yangyang, et al.
Veröffentlicht: (2026)
EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI
von: Kasu, Sai Kartheek Reddy
Veröffentlicht: (2025)
von: Kasu, Sai Kartheek Reddy
Veröffentlicht: (2025)
Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests
von: Noever, David, et al.
Veröffentlicht: (2025)
von: Noever, David, et al.
Veröffentlicht: (2025)
CRAVE: A Conflicting Reasoning Approach for Explainable Claim Verification Using LLMs
von: Zheng, Yingming, et al.
Veröffentlicht: (2025)
von: Zheng, Yingming, et al.
Veröffentlicht: (2025)
When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
von: Qian, Lingfei, et al.
Veröffentlicht: (2025)
von: Qian, Lingfei, et al.
Veröffentlicht: (2025)
MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs
von: Alyafeai, Zaid, et al.
Veröffentlicht: (2025)
von: Alyafeai, Zaid, et al.
Veröffentlicht: (2025)
Peerispect: Claim Verification in Scientific Peer Reviews
von: Ghorbanpour, Ali, et al.
Veröffentlicht: (2026)
von: Ghorbanpour, Ali, et al.
Veröffentlicht: (2026)
VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning
von: Jiang, Zhihuan, et al.
Veröffentlicht: (2024)
von: Jiang, Zhihuan, et al.
Veröffentlicht: (2024)
AI Can Learn Scientific Taste
von: Tong, Jingqi, et al.
Veröffentlicht: (2026)
von: Tong, Jingqi, et al.
Veröffentlicht: (2026)
Can LLMs Reason in the Wild with Programs?
von: Yang, Yuan, et al.
Veröffentlicht: (2024)
von: Yang, Yuan, et al.
Veröffentlicht: (2024)
MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
von: Tian, Yang, et al.
Veröffentlicht: (2025)
von: Tian, Yang, et al.
Veröffentlicht: (2025)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
von: Xu, Zhijian, et al.
Veröffentlicht: (2025)
von: Xu, Zhijian, et al.
Veröffentlicht: (2025)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2024)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
von: Yang, Eddie, et al.
Veröffentlicht: (2026)
von: Yang, Eddie, et al.
Veröffentlicht: (2026)
PreScience: A Benchmark for Forecasting Scientific Contributions
von: Ajith, Anirudh, et al.
Veröffentlicht: (2026)
von: Ajith, Anirudh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2024) -
LLM-Generated Black-box Explanations Can Be Adversarially Helpful
von: Ajwani, Rohan, et al.
Veröffentlicht: (2024) -
Truth Neurons
von: Li, Haohang, et al.
Veröffentlicht: (2025) -
VERBA: Verbalizing Model Differences Using Large Language Models
von: Doda, Shravan, et al.
Veröffentlicht: (2025) -
Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2025)