Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tschisgale, Paul, Maus, Holger, Kieser, Fabian, Kroehs, Ben, Petersen, Stefan, Wulff, Peter |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Developing and Evaluating a Large Language Model-Based Automated Feedback System Grounded in Evidence-Centered Design for Supporting Physics Problem Solving
von: Maus, Holger, et al.
Veröffentlicht: (2025)
von: Maus, Holger, et al.
Veröffentlicht: (2025)
Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
von: Tschisgale, Paul, et al.
Veröffentlicht: (2026)
von: Tschisgale, Paul, et al.
Veröffentlicht: (2026)
Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts
von: Bleckmann, Tom, et al.
Veröffentlicht: (2025)
von: Bleckmann, Tom, et al.
Veröffentlicht: (2025)
PromptCoT: Synthesizing Olympiad-level Problems for Mathematical Reasoning in Large Language Models
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
Meteorological observations during POLARSTERN cruise PS122/1
von: Schmithüsen, Holger, et al.
Veröffentlicht: (2021)
von: Schmithüsen, Holger, et al.
Veröffentlicht: (2021)
Radiosonde raw data measured during POLARSTERN cruise PS134, links to files
von: Schmithüsen, Holger, et al.
Veröffentlicht: (2026)
von: Schmithüsen, Holger, et al.
Veröffentlicht: (2026)
Meteorological observations during POLARSTERN cruise PS134
von: Schmithüsen, Holger, et al.
Veröffentlicht: (2025)
von: Schmithüsen, Holger, et al.
Veröffentlicht: (2025)
Radiosonde measurements during POLARSTERN cruise PS134
von: Schmithüsen, Holger, et al.
Veröffentlicht: (2026)
von: Schmithüsen, Holger, et al.
Veröffentlicht: (2026)
CrawfordGribben and John W.Tweeddale, eds. T&T Clark Handbook of John Owen. London: T&T Clark, 2022, 571pp. $39.95
von: Ty Kieser
Veröffentlicht: (2025)
von: Ty Kieser
Veröffentlicht: (2025)
Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
von: Gao, Songyang, et al.
Veröffentlicht: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
von: Cui, Yiming, et al.
Veröffentlicht: (2025)
von: Cui, Yiming, et al.
Veröffentlicht: (2025)
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
von: He, Chaoqun, et al.
Veröffentlicht: (2024)
von: He, Chaoqun, et al.
Veröffentlicht: (2024)
ChatQA: Surpassing GPT-4 on Conversational QA and RAG
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
RIMO: An Easy-to-Evaluate, Hard-to-Solve Olympiad Benchmark for Advanced Mathematical Reasoning
von: Chen, Ziye, et al.
Veröffentlicht: (2025)
von: Chen, Ziye, et al.
Veröffentlicht: (2025)
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
von: Sun, Haoxiang, et al.
Veröffentlicht: (2025)
von: Sun, Haoxiang, et al.
Veröffentlicht: (2025)
P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads
von: Luo, Yun, et al.
Veröffentlicht: (2026)
von: Luo, Yun, et al.
Veröffentlicht: (2026)
Comments on “Sample Size Adaptation Designs and Efficiency Comparison With Group Sequential Designs”
von: Meinhard Kieser, et al.
Veröffentlicht: (2025)
von: Meinhard Kieser, et al.
Veröffentlicht: (2025)
Artificial Intelligence in Physical Therapy Education: Evaluating Clinical Reasoning Performance in Musculoskeletal Care Using ChatGPT
von: Jie Hao, et al.
Veröffentlicht: (2025)
von: Jie Hao, et al.
Veröffentlicht: (2025)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2025)
von: Mahdavi, Sadegh, et al.
Veröffentlicht: (2025)
Unreflected Acceptance -- Investigating the Negative Consequences of ChatGPT-Assisted Problem Solving in Physics Education
von: Krupp, Lars, et al.
Veröffentlicht: (2023)
von: Krupp, Lars, et al.
Veröffentlicht: (2023)
Proving Olympiad Inequalities by Synergizing LLMs and Symbolic Reasoning
von: Li, Zenan, et al.
Veröffentlicht: (2025)
von: Li, Zenan, et al.
Veröffentlicht: (2025)
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
von: Zhu, Yaoming, et al.
Veröffentlicht: (2025)
von: Zhu, Yaoming, et al.
Veröffentlicht: (2025)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
von: Zou, Kaijian, et al.
Veröffentlicht: (2025)
von: Zou, Kaijian, et al.
Veröffentlicht: (2025)
Linguistics Olympiad
von: Neacșu, Vlad A.
Veröffentlicht: (2024)
von: Neacșu, Vlad A.
Veröffentlicht: (2024)
Proving Olympiad Algebraic Inequalities without Human Demonstrations
von: Wei, Chenrui, et al.
Veröffentlicht: (2024)
von: Wei, Chenrui, et al.
Veröffentlicht: (2024)
Trigonometry and Analytic Tools in Olympiad Geometry Problems, Part I
von: Lignos, Orestis
Veröffentlicht: (2023)
von: Lignos, Orestis
Veröffentlicht: (2023)
Overview of AI Grading of Physics Olympiad Exams
von: McGinness, Lachlan
Veröffentlicht: (2025)
von: McGinness, Lachlan
Veröffentlicht: (2025)
Mastering Olympiad-Level Physics with Artificial Intelligence
von: Jian, Dong-Shan, et al.
Veröffentlicht: (2025)
von: Jian, Dong-Shan, et al.
Veröffentlicht: (2025)
CogAtom: From Cognitive Atoms to Olympiad-level Mathematical Reasoning in Large Language Models
von: Chen, Zhuofan, et al.
Veröffentlicht: (2025)
von: Chen, Zhuofan, et al.
Veröffentlicht: (2025)
Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2026)
von: Prabhudesai, Mihir, et al.
Veröffentlicht: (2026)
FormalGeo: An Extensible Formalized Framework for Olympiad Geometric Problem Solving
von: Zhang, Xiaokai, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaokai, et al.
Veröffentlicht: (2023)
Large Language Models Achieve Gold Medal Performance at the International Olympiad on Astronomy & Astrophysics (IOAA)
von: Pinheiro, Lucas Carrit Delgado, et al.
Veröffentlicht: (2025)
von: Pinheiro, Lucas Carrit Delgado, et al.
Veröffentlicht: (2025)
P1: Mastering Physics Olympiads with Reinforcement Learning
von: Chen, Jiacheng, et al.
Veröffentlicht: (2025)
von: Chen, Jiacheng, et al.
Veröffentlicht: (2025)
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
Brains vs. Bytes: Evaluating LLM Proficiency in Olympiad Mathematics
von: Mahdavi, Hamed, et al.
Veröffentlicht: (2025)
von: Mahdavi, Hamed, et al.
Veröffentlicht: (2025)
Assessing UML Diagrams by GPT: Implications for Education
von: Wang, Chong, et al.
Veröffentlicht: (2024)
von: Wang, Chong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Developing and Evaluating a Large Language Model-Based Automated Feedback System Grounded in Evidence-Centered Design for Supporting Physics Problem Solving
von: Maus, Holger, et al.
Veröffentlicht: (2025) -
Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research
von: Tschisgale, Paul, et al.
Veröffentlicht: (2026) -
Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts
von: Bleckmann, Tom, et al.
Veröffentlicht: (2025) -
PromptCoT: Synthesizing Olympiad-level Problems for Mathematical Reasoning in Large Language Models
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025) -
Meteorological observations during POLARSTERN cruise PS122/1
von: Schmithüsen, Holger, et al.
Veröffentlicht: (2021)