SciIF: Benchmarking Scientific Instruction Following Towards Rigorous Scientific Intelligence
Fuente:
arXiv
Guardado en:
| Autores principales: | Su, Encheng, Wu, Jianyu, Tang, Chen, Wang, Lintao, Li, Pengze, Wang, Aoran, Zhang, Jinouwen, Wang, Yizhou, Meng, Yuan, Ma, Xinzhu, Tang, Shixiang, Li, Houqiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
por: Su, Encheng, et al.
Publicado: (2026)
por: Su, Encheng, et al.
Publicado: (2026)
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
por: Wang, Lintao, et al.
Publicado: (2025)
por: Wang, Lintao, et al.
Publicado: (2025)
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
por: Wang, Yizhou, et al.
Publicado: (2025)
por: Wang, Yizhou, et al.
Publicado: (2025)
SciDataCopilot: An Agentic Data Preparation Framework for AGI-driven Scientific Discovery
por: Rao, Jiyong, et al.
Publicado: (2026)
por: Rao, Jiyong, et al.
Publicado: (2026)
SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
por: Yu, Wenhan, et al.
Publicado: (2025)
por: Yu, Wenhan, et al.
Publicado: (2025)
An Intelligent Innovation Dataset on Scientific Research Outcomes
por: Wu, Xinran, et al.
Publicado: (2024)
por: Wu, Xinran, et al.
Publicado: (2024)
Charting Empirical Laws for LLM Fine-Tuning in Scientific Multi-Discipline Learning
por: Wang, Lintao, et al.
Publicado: (2026)
por: Wang, Lintao, et al.
Publicado: (2026)
Benchmarking Table Extraction from Heterogeneous Scientific Extraction Documents
por: Soric, Marijan, et al.
Publicado: (2025)
por: Soric, Marijan, et al.
Publicado: (2025)
Automated Extraction of Mechanical Constitutive Models from Scientific Literature using Large Language Models: Applications in Cultural Heritage Conservation
por: Hu, Rui, et al.
Publicado: (2026)
por: Hu, Rui, et al.
Publicado: (2026)
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
por: Wang, Yiheng, et al.
Publicado: (2025)
por: Wang, Yiheng, et al.
Publicado: (2025)
The State of Scientific Poster Sharing and Reuse
por: Gasimova, Aydan, et al.
Publicado: (2026)
por: Gasimova, Aydan, et al.
Publicado: (2026)
GenIE - Simulator-Driven Iterative Data Exploration for Scientific Discovery
por: Colaco, Ashwin Gerard, et al.
Publicado: (2025)
por: Colaco, Ashwin Gerard, et al.
Publicado: (2025)
LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
por: Li, Rui, et al.
Publicado: (2025)
por: Li, Rui, et al.
Publicado: (2025)
OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads
por: Hwang, Soon, et al.
Publicado: (2025)
por: Hwang, Soon, et al.
Publicado: (2025)
Manifesto for Scientifically Sound Artificial Intelligence Towards an Artificial Intelligence Serving Scientific Rigor
por: Febba, Michel
Publicado: (2025)
por: Febba, Michel
Publicado: (2025)
Enabling Homomorphic Analytical Operations on Compressed Scientific Data with Multi-stage Decompression
por: Wu, Xuan, et al.
Publicado: (2026)
por: Wu, Xuan, et al.
Publicado: (2026)
A Declarative Query Language for Scientific Machine Learning
por: Jamil, Hasan M
Publicado: (2024)
por: Jamil, Hasan M
Publicado: (2024)
NL2SQL-BUGs: A Benchmark for Detecting Semantic Errors in NL2SQL Translation
por: Liu, Xinyu, et al.
Publicado: (2025)
por: Liu, Xinyu, et al.
Publicado: (2025)
From Tokens to Materials: Leveraging Language Models for Scientific Discovery
por: Wan, Yuwei, et al.
Publicado: (2024)
por: Wan, Yuwei, et al.
Publicado: (2024)
Implementing a Scalable, Redeployable and Multitiered Repository for FAIR and Secure Scientific Data Sharing: The BIG-MAP Archive
por: Granata, Valeria, et al.
Publicado: (2025)
por: Granata, Valeria, et al.
Publicado: (2025)
LCP: Enhancing Scientific Data Management with Lossy Compression for Particles
por: Zhang, Longtao, et al.
Publicado: (2024)
por: Zhang, Longtao, et al.
Publicado: (2024)
Applying Process Mining on Scientific Workflows: a Case Study on High Performance Computing Data
por: Sadeghibogar, Zahra, et al.
Publicado: (2023)
por: Sadeghibogar, Zahra, et al.
Publicado: (2023)
Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
por: Tian, Jiaming, et al.
Publicado: (2025)
por: Tian, Jiaming, et al.
Publicado: (2025)
Ranking Business and Economics Journals in South America Using the Scientific Electronic Library Online (SciELO)
por: Alexander, Jennifer K., et al.
Publicado: (2012)
por: Alexander, Jennifer K., et al.
Publicado: (2012)
ChatBI: Towards Natural Language to Complex Business Intelligence SQL
por: Lian, Jinqing, et al.
Publicado: (2024)
por: Lian, Jinqing, et al.
Publicado: (2024)
GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian Optimization
por: Lao, Jiale, et al.
Publicado: (2023)
por: Lao, Jiale, et al.
Publicado: (2023)
CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
por: Deng, Han, et al.
Publicado: (2025)
por: Deng, Han, et al.
Publicado: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
por: Chen, Haotian, et al.
Publicado: (2025)
por: Chen, Haotian, et al.
Publicado: (2025)
CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
por: Wu, Jianyu, et al.
Publicado: (2025)
por: Wu, Jianyu, et al.
Publicado: (2025)
LLM-FK: Multi-Agent LLM Reasoning for Foreign Key Detection in Large-Scale Complex Databases
por: Tang, Zijian, et al.
Publicado: (2026)
por: Tang, Zijian, et al.
Publicado: (2026)
MCI-SQL: Text-to-SQL with Metadata-Complete Context and Intermediate Correction
por: Wang, Qin, et al.
Publicado: (2026)
por: Wang, Qin, et al.
Publicado: (2026)
Development of Data Evaluation Benchmark for Data Wrangling Recommendation System
por: Wang, Yuqing, et al.
Publicado: (2024)
por: Wang, Yuqing, et al.
Publicado: (2024)
DB-GPT-Hub: Towards Open Benchmarking Text-to-SQL Empowered by Large Language Models
por: Zhou, Fan, et al.
Publicado: (2024)
por: Zhou, Fan, et al.
Publicado: (2024)
Towards Temporal Knowledge Graph Alignment in the Wild
por: Zhao, Runhao, et al.
Publicado: (2025)
por: Zhao, Runhao, et al.
Publicado: (2025)
ProvDeploy: Provenance-oriented Containerization of High Performance Computing Scientific Workflows
por: Kunstmann, Liliane, et al.
Publicado: (2024)
por: Kunstmann, Liliane, et al.
Publicado: (2024)
DTBench: A Synthetic Benchmark for Document-to-Table Extraction
por: Guo, Yuxiang, et al.
Publicado: (2026)
por: Guo, Yuxiang, et al.
Publicado: (2026)
The Dawn of Natural Language to SQL: Are We Fully Ready?
por: Li, Boyan, et al.
Publicado: (2024)
por: Li, Boyan, et al.
Publicado: (2024)
Data Agents: Levels, State of the Art, and Open Problems
por: Luo, Yuyu, et al.
Publicado: (2026)
por: Luo, Yuyu, et al.
Publicado: (2026)
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
por: Cai, Hengxing, et al.
Publicado: (2024)
por: Cai, Hengxing, et al.
Publicado: (2024)
Relational Database Augmented Large Language Model
por: Qin, Zongyue, et al.
Publicado: (2024)
por: Qin, Zongyue, et al.
Publicado: (2024)
Ejemplares similares
-
S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
por: Su, Encheng, et al.
Publicado: (2026) -
PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
por: Wang, Lintao, et al.
Publicado: (2025) -
SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
por: Wang, Yizhou, et al.
Publicado: (2025) -
SciDataCopilot: An Agentic Data Preparation Framework for AGI-driven Scientific Discovery
por: Rao, Jiyong, et al.
Publicado: (2026) -
SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
por: Yu, Wenhan, et al.
Publicado: (2025)