FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Miles, Lin, Robi, Hu, Kat, Jiao, Joy, Chowdhury, Neil, Chang, Ethan, Patwardhan, Tejal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
by: Patwardhan, Tejal, et al.
Published: (2025)
by: Patwardhan, Tejal, et al.
Published: (2025)
Does Distributed Training Undermine Compute Governance?
by: Rahman, Robi
Published: (2026)
by: Rahman, Robi
Published: (2026)
Trends in AI Supercomputers
by: Pilz, Konstantin F., et al.
Published: (2025)
by: Pilz, Konstantin F., et al.
Published: (2025)
Evaluating AI Providers' Frontier Safety Frameworks
by: Stelling, Lily, et al.
Published: (2025)
by: Stelling, Lily, et al.
Published: (2025)
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
by: Miserendino, Samuel, et al.
Published: (2025)
by: Miserendino, Samuel, et al.
Published: (2025)
The rising costs of training frontier AI models
by: Cottier, Ben, et al.
Published: (2024)
by: Cottier, Ben, et al.
Published: (2024)
Computational Thinking with Computer Vision: Developing AI Competency in an Introductory Computer Science Course
by: Chowdhury, Tahiya
Published: (2025)
by: Chowdhury, Tahiya
Published: (2025)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
A Possibility Frontier Approach to Diverse Talent Selection
by: Natarajan, Neil, et al.
Published: (2025)
by: Natarajan, Neil, et al.
Published: (2025)
Expanding External Access To Frontier AI Models For Dangerous Capability Evaluations
by: Charnock, Jacob, et al.
Published: (2026)
by: Charnock, Jacob, et al.
Published: (2026)
WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
by: Maurya, Sneha, et al.
Published: (2026)
by: Maurya, Sneha, et al.
Published: (2026)
Critically Engaged Pragmatism: A Scientific Norm and Social, Pragmatist Epistemology for AI Science Evaluation Tools
by: Lee, Carole J.
Published: (2026)
by: Lee, Carole J.
Published: (2026)
Sabotage Evaluations for Frontier Models
by: Benton, Joe, et al.
Published: (2024)
by: Benton, Joe, et al.
Published: (2024)
Howzat? Appealing to Expert Judgement for Evaluating Human and AI Next-Step Hints for Novice Programmers
by: Brown, Neil C. C., et al.
Published: (2024)
by: Brown, Neil C. C., et al.
Published: (2024)
PaperBench: Evaluating AI's Ability to Replicate AI Research
by: Starace, Giulio, et al.
Published: (2025)
by: Starace, Giulio, et al.
Published: (2025)
Evaluating the AI-Lab Intervention: Impact on Student Perception and Use of Generative AI in Early Undergraduate Computer Science Courses
by: Dickey, Ethan, et al.
Published: (2025)
by: Dickey, Ethan, et al.
Published: (2025)
The CitizenQuery Benchmark: A Novel Dataset and Evaluation Pipeline for Measuring LLM Performance in Citizen Query Tasks
by: Majithia, Neil, et al.
Published: (2026)
by: Majithia, Neil, et al.
Published: (2026)
The California Report on Frontier AI Policy
by: Bommasani, Rishi, et al.
Published: (2025)
by: Bommasani, Rishi, et al.
Published: (2025)
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
by: Kochmar, Ekaterina, et al.
Published: (2025)
by: Kochmar, Ekaterina, et al.
Published: (2025)
Frontier AI Ethics: Anticipating and Evaluating the Societal Impacts of Language Model Agents
by: Lazar, Seth
Published: (2024)
by: Lazar, Seth
Published: (2024)
Assurance of Frontier AI Built for National Security
by: Pistillo, Matteo, et al.
Published: (2025)
by: Pistillo, Matteo, et al.
Published: (2025)
Preliminary Analysis of Construction Work Zone on Roadways in Florida by Crash Severity
by: Deslouches, Tatiana, et al.
Published: (2025)
by: Deslouches, Tatiana, et al.
Published: (2025)
Bridging Gaps Between Student and Expert Evaluations of AI-Generated Programming Hints
by: Phung, Tung, et al.
Published: (2025)
by: Phung, Tung, et al.
Published: (2025)
Evaluating Performance Consistency in Competitive Programming: Educational Implications and Contest Design Insights
by: Luo, Zhongtang, et al.
Published: (2025)
by: Luo, Zhongtang, et al.
Published: (2025)
Responsible Reporting for Frontier AI Development
by: Kolt, Noam, et al.
Published: (2024)
by: Kolt, Noam, et al.
Published: (2024)
Governing AI Beyond the Pretraining Frontier
by: Caputo, Nicholas A.
Published: (2025)
by: Caputo, Nicholas A.
Published: (2025)
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024)
by: Wallach, Hanna, et al.
Published: (2024)
Frontier AI developers need an internal audit function
by: Schuett, Jonas
Published: (2023)
by: Schuett, Jonas
Published: (2023)
The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety
by: Fort, Kristina
Published: (2024)
by: Fort, Kristina
Published: (2024)
Science Literacy: Generative AI as Enabler of Coherence in the Teaching, Learning, and Assessment of Scientific Knowledge and Reasoning
by: Zhai, Xiaoming, et al.
Published: (2026)
by: Zhai, Xiaoming, et al.
Published: (2026)
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
by: Brundage, Miles, et al.
Published: (2026)
by: Brundage, Miles, et al.
Published: (2026)
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
by: Gringras, David, et al.
Published: (2026)
by: Gringras, David, et al.
Published: (2026)
Collective Predictive Coding as Model of Science: Formalizing Scientific Activities Towards Generative Science
by: Taniguchi, Tadahiro, et al.
Published: (2024)
by: Taniguchi, Tadahiro, et al.
Published: (2024)
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
by: Tong, Haibo, et al.
Published: (2026)
by: Tong, Haibo, et al.
Published: (2026)
Insights for an AI Whistleblower Office from 30 Case Studies
by: Beri, Ethan, et al.
Published: (2026)
by: Beri, Ethan, et al.
Published: (2026)
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2025)
by: Wallach, Hanna, et al.
Published: (2025)
Certified Safe: A Schematic for Approval Regulation of Frontier AI
by: Salvador, Cole
Published: (2024)
by: Salvador, Cole
Published: (2024)
Enabling Frontier Lab Collaboration to Mitigate AI Safety Risks
by: Felstead, Nicholas
Published: (2025)
by: Felstead, Nicholas
Published: (2025)
From Principles to Rules: A Regulatory Approach for Frontier AI
by: Schuett, Jonas, et al.
Published: (2024)
by: Schuett, Jonas, et al.
Published: (2024)
When Does Regulation by Insurance Work? The Case of Frontier AI
by: Trout, Cristian
Published: (2025)
by: Trout, Cristian
Published: (2025)
Similar Items
-
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
by: Patwardhan, Tejal, et al.
Published: (2025) -
Does Distributed Training Undermine Compute Governance?
by: Rahman, Robi
Published: (2026) -
Trends in AI Supercomputers
by: Pilz, Konstantin F., et al.
Published: (2025) -
Evaluating AI Providers' Frontier Safety Frameworks
by: Stelling, Lily, et al.
Published: (2025) -
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
by: Miserendino, Samuel, et al.
Published: (2025)