SciTaRC: Benchmarking QA on Scientific Tabular Data that Requires Language Reasoning and Complex Computation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hexuan, Ren, Yaxuan, Bommireddypalli, Srikar, Chen, Shuxian, Prabhudesai, Adarsh, Zhou, Rongkun, Baral, Elina, Koehn, Philipp |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CRAFT: Training-Free Cascaded Retrieval for Tabular QA
von: Singh, Adarsh, et al.
Veröffentlicht: (2025)
von: Singh, Adarsh, et al.
Veröffentlicht: (2025)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2026)
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2026)
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
von: Li, Yuyi, et al.
Veröffentlicht: (2025)
von: Li, Yuyi, et al.
Veröffentlicht: (2025)
SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
von: Deng, Andong, et al.
Veröffentlicht: (2025)
von: Deng, Andong, et al.
Veröffentlicht: (2025)
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
von: Arazi, Alan, et al.
Veröffentlicht: (2026)
von: Arazi, Alan, et al.
Veröffentlicht: (2026)
Text Style Transfer with Parameter-efficient LLM Finetuning and Round-trip Translation
von: Liu, Ruoxi, et al.
Veröffentlicht: (2026)
von: Liu, Ruoxi, et al.
Veröffentlicht: (2026)
Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents
von: Meng, Chutong, et al.
Veröffentlicht: (2025)
von: Meng, Chutong, et al.
Veröffentlicht: (2025)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
von: Lu, Taiming, et al.
Veröffentlicht: (2024)
How Robust are the Tabular QA Models for Scientific Tables? A Study using Customized Dataset
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
von: Su, Encheng, et al.
Veröffentlicht: (2026)
von: Su, Encheng, et al.
Veröffentlicht: (2026)
Geometry Matters: Benchmarking Scientific ML Approaches for Flow Prediction around Complex Geometries
von: Rabeh, Ali, et al.
Veröffentlicht: (2024)
von: Rabeh, Ali, et al.
Veröffentlicht: (2024)
Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon
von: Gallego, Víctor
Veröffentlicht: (2026)
von: Gallego, Víctor
Veröffentlicht: (2026)
SciMDR: Advancing Scientific Multimodal Document Reasoning
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
von: Wu, Siwei, et al.
Veröffentlicht: (2024)
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
von: Cai, Hengxing, et al.
Veröffentlicht: (2024)
von: Cai, Hengxing, et al.
Veröffentlicht: (2024)
SciEvent: Benchmarking Multi-domain Scientific Event Extraction
von: Dong, Bofu, et al.
Veröffentlicht: (2025)
von: Dong, Bofu, et al.
Veröffentlicht: (2025)
SciAgent: Tool-augmented Language Models for Scientific Reasoning
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
von: Liu, Tengxiao, et al.
Veröffentlicht: (2026)
von: Liu, Tengxiao, et al.
Veröffentlicht: (2026)
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
Quantifying Reproducibility Gaps in Publicly Available Plant Nuclear Bioimaging Datasets: The Reproducibility Risk Assessment Framework (RRAF)
von: Joardar, Sudipta, et al.
Veröffentlicht: (2026)
von: Joardar, Sudipta, et al.
Veröffentlicht: (2026)
Measuring Visual Understanding in Telecom domain: Performance Metrics for Image-to-UML conversion using VLMs
von: Ranjani, HG, et al.
Veröffentlicht: (2025)
von: Ranjani, HG, et al.
Veröffentlicht: (2025)
Dynamics of Dissociative Electron Attachment to Aliphatic Thiols
von: Das, Sukanta, et al.
Veröffentlicht: (2023)
von: Das, Sukanta, et al.
Veröffentlicht: (2023)
Breakthrough Asymmetries across Disciplines and Countries: A Network approach to Structural Complexity of Scientific Progress
von: Raghuvanshi, Adarsh, et al.
Veröffentlicht: (2025)
von: Raghuvanshi, Adarsh, et al.
Veröffentlicht: (2025)
SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation
von: Roberts, Jonathan, et al.
Veröffentlicht: (2024)
von: Roberts, Jonathan, et al.
Veröffentlicht: (2024)
GeoRC: A Benchmark for Geolocation Reasoning Chains
von: Talreja, Mohit, et al.
Veröffentlicht: (2026)
von: Talreja, Mohit, et al.
Veröffentlicht: (2026)
RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
von: Satriani, Dario, et al.
Veröffentlicht: (2025)
von: Satriani, Dario, et al.
Veröffentlicht: (2025)
SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning
von: Zheng, Tianshi, et al.
Veröffentlicht: (2026)
von: Zheng, Tianshi, et al.
Veröffentlicht: (2026)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
von: Burgess, James, et al.
Veröffentlicht: (2026)
von: Burgess, James, et al.
Veröffentlicht: (2026)
Pointer-Generator Networks for Low-Resource Machine Translation: Don't Copy That!
von: Bafna, Niyati, et al.
Veröffentlicht: (2024)
von: Bafna, Niyati, et al.
Veröffentlicht: (2024)
Recovering document annotations for sentence-level bitext
von: Wicks, Rachel, et al.
Veröffentlicht: (2024)
von: Wicks, Rachel, et al.
Veröffentlicht: (2024)
Justinian und die Armee des frühen Byzanz
von: Koehn, Clemens
Veröffentlicht: (2021)
von: Koehn, Clemens
Veröffentlicht: (2021)
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
von: Elkoussy, Laïla, et al.
Veröffentlicht: (2026)
von: Elkoussy, Laïla, et al.
Veröffentlicht: (2026)
SciFigDetect: A Benchmark for AI-Generated Scientific Figure Detection
von: Hu, You, et al.
Veröffentlicht: (2026)
von: Hu, You, et al.
Veröffentlicht: (2026)
SciDesignBench: Benchmarking and Improving Language Models for Scientific Inverse Design
von: van Dijk, David, et al.
Veröffentlicht: (2026)
von: van Dijk, David, et al.
Veröffentlicht: (2026)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
von: Maheshwary, Rishabh, et al.
Veröffentlicht: (2025)
von: Maheshwary, Rishabh, et al.
Veröffentlicht: (2025)
Scientific QA System with Verifiable Answers
von: Ljajić, Adela, et al.
Veröffentlicht: (2024)
von: Ljajić, Adela, et al.
Veröffentlicht: (2024)
SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
von: Yu, Wenhan, et al.
Veröffentlicht: (2025)
von: Yu, Wenhan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CRAFT: Training-Free Cascaded Retrieval for Tabular QA
von: Singh, Adarsh, et al.
Veröffentlicht: (2025) -
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024) -
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2026) -
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
von: Li, Yuyi, et al.
Veröffentlicht: (2025) -
SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation
von: Guo, Longteng, et al.
Veröffentlicht: (2026)