Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science
Fuente:
arXiv
Guardado en:
| Autores principales: | Jansen, Peter, Hassan, Samiah, Wang, Ruoyao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials
por: Lv, Taoyuze, et al.
Publicado: (2025)
por: Lv, Taoyuze, et al.
Publicado: (2025)
HealthFC: Verifying Health Claims with Evidence-Based Medical Fact-Checking
por: Vladika, Juraj, et al.
Publicado: (2023)
por: Vladika, Juraj, et al.
Publicado: (2023)
Evaluating the Performance and Robustness of LLMs in Materials Science Q&A and Property Predictions
por: Wang, Hongchen, et al.
Publicado: (2024)
por: Wang, Hongchen, et al.
Publicado: (2024)
VeriFact: Verifying Facts in LLM-Generated Clinical Text with Electronic Health Records
por: Chung, Philip, et al.
Publicado: (2025)
por: Chung, Philip, et al.
Publicado: (2025)
Flexible, Model-Agnostic Method for Materials Data Extraction from Text Using General Purpose Language Models
por: Polak, Maciej P., et al.
Publicado: (2023)
por: Polak, Maciej P., et al.
Publicado: (2023)
LLMatDesign: Autonomous Materials Discovery with Large Language Models
por: Jia, Shuyi, et al.
Publicado: (2024)
por: Jia, Shuyi, et al.
Publicado: (2024)
MatExpert: Decomposing Materials Discovery by Mimicking Human Experts
por: Ding, Qianggang, et al.
Publicado: (2024)
por: Ding, Qianggang, et al.
Publicado: (2024)
Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
por: Hyun, Lee, et al.
Publicado: (2025)
por: Hyun, Lee, et al.
Publicado: (2025)
LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation
por: Chiang, Yuan, et al.
Publicado: (2024)
por: Chiang, Yuan, et al.
Publicado: (2024)
Robust Claim Verification Through Fact Detection
por: Jafari, Nazanin, et al.
Publicado: (2024)
por: Jafari, Nazanin, et al.
Publicado: (2024)
Language-Native Materials Processing Design by Lightly Structured Text Database and Reasoning Large Language Model
por: Liu, Yuze, et al.
Publicado: (2025)
por: Liu, Yuze, et al.
Publicado: (2025)
BeamPERL: Parameter-Efficient RL with Verifiable Rewards Specializes Compact LLMs for Structured Beam Mechanics Reasoning
por: Hage, Tarjei Paule, et al.
Publicado: (2026)
por: Hage, Tarjei Paule, et al.
Publicado: (2026)
Can Language Models Serve as Text-Based World Simulators?
por: Wang, Ruoyao, et al.
Publicado: (2024)
por: Wang, Ruoyao, et al.
Publicado: (2024)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
por: Seo, Wooseok, et al.
Publicado: (2025)
por: Seo, Wooseok, et al.
Publicado: (2025)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
por: Sundriyal, Megha, et al.
Publicado: (2023)
por: Sundriyal, Megha, et al.
Publicado: (2023)
Piecing It All Together: Verifying Multi-Hop Multimodal Claims
por: Wang, Haoran, et al.
Publicado: (2024)
por: Wang, Haoran, et al.
Publicado: (2024)
Are LLMs Ready for Real-World Materials Discovery?
por: Miret, Santiago, et al.
Publicado: (2024)
por: Miret, Santiago, et al.
Publicado: (2024)
Step-by-Step Fact Verification System for Medical Claims with Explainable Reasoning
por: Vladika, Juraj, et al.
Publicado: (2025)
por: Vladika, Juraj, et al.
Publicado: (2025)
ClaimCheck: Real-Time Fact-Checking with Small Language Models
por: Putta, Akshith Reddy, et al.
Publicado: (2025)
por: Putta, Akshith Reddy, et al.
Publicado: (2025)
Ta'keed: The First Generative Fact-Checking System for Arabic Claims
por: Althabiti, Saud, et al.
Publicado: (2024)
por: Althabiti, Saud, et al.
Publicado: (2024)
Automated Extraction of Material Properties using LLM-based AI Agents
por: Ghosh, Subham, et al.
Publicado: (2025)
por: Ghosh, Subham, et al.
Publicado: (2025)
ClaimIQ at CheckThat! 2025: Comparing Prompted and Fine-Tuned Language Models for Verifying Numerical Claims
por: Anik, Anirban Saha, et al.
Publicado: (2025)
por: Anik, Anirban Saha, et al.
Publicado: (2025)
Multimodal Claim Extraction for Fact-Checking
por: Teo, Joycelyn, et al.
Publicado: (2026)
por: Teo, Joycelyn, et al.
Publicado: (2026)
BiDeV: Bilateral Defusing Verification for Complex Claim Fact-Checking
por: Liu, Yuxuan, et al.
Publicado: (2025)
por: Liu, Yuxuan, et al.
Publicado: (2025)
Multimodal Fact-Level Attribution for Verifiable Reasoning
por: Wan, David, et al.
Publicado: (2026)
por: Wan, David, et al.
Publicado: (2026)
Learning to Verify Summary Facts with Fine-Grained LLM Feedback
por: Oh, Jihwan, et al.
Publicado: (2024)
por: Oh, Jihwan, et al.
Publicado: (2024)
Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence
por: Wang, Fiona Y., et al.
Publicado: (2026)
por: Wang, Fiona Y., et al.
Publicado: (2026)
Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy
por: Roy, Aritra, et al.
Publicado: (2026)
por: Roy, Aritra, et al.
Publicado: (2026)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
por: Liu, Hongjun, et al.
Publicado: (2025)
por: Liu, Hongjun, et al.
Publicado: (2025)
Reshaping MOFs text mining with a dynamic multi-agents framework of large language model
por: Lin, Zuhong, et al.
Publicado: (2025)
por: Lin, Zuhong, et al.
Publicado: (2025)
Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities
por: Lu, Wei, et al.
Publicado: (2024)
por: Lu, Wei, et al.
Publicado: (2024)
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
por: Javaji, Shashidhar Reddy, et al.
Publicado: (2025)
por: Javaji, Shashidhar Reddy, et al.
Publicado: (2025)
CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents
por: Jansen, Peter, et al.
Publicado: (2025)
por: Jansen, Peter, et al.
Publicado: (2025)
Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance?
por: Hu, Qisheng, et al.
Publicado: (2024)
por: Hu, Qisheng, et al.
Publicado: (2024)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
por: Yang, Shuo, et al.
Publicado: (2025)
por: Yang, Shuo, et al.
Publicado: (2025)
Generating Literature-Driven Scientific Theories at Scale
por: Jansen, Peter, et al.
Publicado: (2026)
por: Jansen, Peter, et al.
Publicado: (2026)
Towards Automated Fact-Checking of Real-World Claims: Exploring Task Formulation and Assessment with LLMs
por: Sahitaj, Premtim, et al.
Publicado: (2025)
por: Sahitaj, Premtim, et al.
Publicado: (2025)
Zero-shot and Few-shot Learning with Instruction-following LLMs for Claim Matching in Automated Fact-checking
por: Pisarevskaya, Dina, et al.
Publicado: (2025)
por: Pisarevskaya, Dina, et al.
Publicado: (2025)
How LLMs Fail to Support Fact-Checking
por: Proma, Adiba Mahbub, et al.
Publicado: (2025)
por: Proma, Adiba Mahbub, et al.
Publicado: (2025)
Generative Chemical Language Models for Energetic Materials Discovery
por: Salij, Andrew, et al.
Publicado: (2026)
por: Salij, Andrew, et al.
Publicado: (2026)
Ejemplares similares
-
AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials
por: Lv, Taoyuze, et al.
Publicado: (2025) -
HealthFC: Verifying Health Claims with Evidence-Based Medical Fact-Checking
por: Vladika, Juraj, et al.
Publicado: (2023) -
Evaluating the Performance and Robustness of LLMs in Materials Science Q&A and Property Predictions
por: Wang, Hongchen, et al.
Publicado: (2024) -
VeriFact: Verifying Facts in LLM-Generated Clinical Text with Electronic Health Records
por: Chung, Philip, et al.
Publicado: (2025) -
Flexible, Model-Agnostic Method for Materials Data Extraction from Text Using General Purpose Language Models
por: Polak, Maciej P., et al.
Publicado: (2023)