SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Chengye, Shen, Yifei, Kuang, Zexi, Cohan, Arman, Zhao, Yilun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
SciMDR: Advancing Scientific Multimodal Document Reasoning
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
von: Wang, Chengye, et al.
Veröffentlicht: (2026)
von: Wang, Chengye, et al.
Veröffentlicht: (2026)
MuSciClaims: Multimodal Scientific Claim Verification
von: Lal, Yash Kumar, et al.
Veröffentlicht: (2025)
von: Lal, Yash Kumar, et al.
Veröffentlicht: (2025)
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2024)
von: Zhao, Yilun, et al.
Veröffentlicht: (2024)
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
von: Ding, Hang, et al.
Veröffentlicht: (2025)
von: Ding, Hang, et al.
Veröffentlicht: (2025)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
SciDQA: A Deep Reading Comprehension Dataset over Scientific Papers
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL
von: Shen, Yifei, et al.
Veröffentlicht: (2026)
von: Shen, Yifei, et al.
Veröffentlicht: (2026)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
von: Xu, Zhijian, et al.
Veröffentlicht: (2025)
von: Xu, Zhijian, et al.
Veröffentlicht: (2025)
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
von: Ho, Xanh, et al.
Veröffentlicht: (2026)
von: Ho, Xanh, et al.
Veröffentlicht: (2026)
SciClaimHunt: A Large Dataset for Evidence-based Scientific Claim Verification
von: Kumar, Sujit, et al.
Veröffentlicht: (2025)
von: Kumar, Sujit, et al.
Veröffentlicht: (2025)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
Table-R1: Inference-Time Scaling for Table Reasoning
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
von: Hu, Tiansheng, et al.
Veröffentlicht: (2025)
von: Hu, Tiansheng, et al.
Veröffentlicht: (2025)
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
Atomic Reasoning for Scientific Table Claim Verification
von: Zhang, Yuji, et al.
Veröffentlicht: (2025)
von: Zhang, Yuji, et al.
Veröffentlicht: (2025)
LimRank: Less is More for Reasoning-Intensive Information Reranking
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
von: Hu, Tiansheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiansheng, et al.
Veröffentlicht: (2026)
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
von: Zhang, Siyue, et al.
Veröffentlicht: (2025)
von: Zhang, Siyue, et al.
Veröffentlicht: (2025)
MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
von: Phanse, Rohan, et al.
Veröffentlicht: (2025)
von: Phanse, Rohan, et al.
Veröffentlicht: (2025)
Z1: Efficient Test-time Scaling with Code
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
AlphaResearch: Accelerating New Algorithm Discovery with Language Models
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
Peerispect: Claim Verification in Scientific Peer Reviews
von: Ghorbanpour, Ali, et al.
Veröffentlicht: (2026)
von: Ghorbanpour, Ali, et al.
Veröffentlicht: (2026)
Step-Back Profiling: Distilling User History for Personalized Scientific Writing
von: Tang, Xiangru, et al.
Veröffentlicht: (2024)
von: Tang, Xiangru, et al.
Veröffentlicht: (2024)
NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
von: Rao, Delip, et al.
Veröffentlicht: (2025)
von: Rao, Delip, et al.
Veröffentlicht: (2025)
IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery
von: Garikaparthi, Aniketh, et al.
Veröffentlicht: (2025)
von: Garikaparthi, Aniketh, et al.
Veröffentlicht: (2025)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature
von: Wadden, David, et al.
Veröffentlicht: (2024)
von: Wadden, David, et al.
Veröffentlicht: (2024)
Unveiling the Spectrum of Data Contamination in Language Models: A Survey from Detection to Remediation
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
von: Zhao, Yilun, et al.
Veröffentlicht: (2025) -
SciMDR: Advancing Scientific Multimodal Document Reasoning
von: Chen, Ziyu, et al.
Veröffentlicht: (2026) -
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024) -
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
von: Wang, Chengye, et al.
Veröffentlicht: (2026) -
MuSciClaims: Multimodal Scientific Claim Verification
von: Lal, Yash Kumar, et al.
Veröffentlicht: (2025)