Saved in:
| Main Authors: | Tschisgale, Paul, Wulff, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.15889 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts
by: Bleckmann, Tom, et al.
Published: (2025)
by: Bleckmann, Tom, et al.
Published: (2025)
Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment
by: Tschisgale, Paul, et al.
Published: (2025)
by: Tschisgale, Paul, et al.
Published: (2025)
Developing and Evaluating a Large Language Model-Based Automated Feedback System Grounded in Evidence-Centered Design for Supporting Physics Problem Solving
by: Maus, Holger, et al.
Published: (2025)
by: Maus, Holger, et al.
Published: (2025)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025)
by: Kuzmanko, Jonathan
Published: (2025)
Metacognitive Myopia in Large Language Models
by: Scholten, Florian, et al.
Published: (2024)
by: Scholten, Florian, et al.
Published: (2024)
The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
by: Dobbins, Nic, et al.
Published: (2025)
by: Dobbins, Nic, et al.
Published: (2025)
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
by: Devunuri, Saipraneeth, et al.
Published: (2024)
by: Devunuri, Saipraneeth, et al.
Published: (2024)
Limits of Large Language Models in Debating Humans
by: Flamino, James, et al.
Published: (2024)
by: Flamino, James, et al.
Published: (2024)
Domain-Shift-Aware Conformal Prediction for Large Language Models
by: Lin, Zhexiao, et al.
Published: (2025)
by: Lin, Zhexiao, et al.
Published: (2025)
Uncertainty-Aware Adaptation of Large Language Models for Protein-Protein Interaction Analysis
by: Jantre, Sanket, et al.
Published: (2025)
by: Jantre, Sanket, et al.
Published: (2025)
Augmented Risk Prediction for the Onset of Alzheimer's Disease from Electronic Health Records with Large Language Models
by: Wang, Jiankun, et al.
Published: (2024)
by: Wang, Jiankun, et al.
Published: (2024)
From Traditional Taggers to LLMs: A Comparative Study of POS Tagging for Medieval Romance Languages
by: Schöffel, Matthias, et al.
Published: (2026)
by: Schöffel, Matthias, et al.
Published: (2026)
Assessing Large Language Models in Mechanical Engineering Education: A Study on Mechanics-Focused Conceptual Understanding
by: Tian, Jie, et al.
Published: (2024)
by: Tian, Jie, et al.
Published: (2024)
Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
by: Hang, Haotian, et al.
Published: (2025)
by: Hang, Haotian, et al.
Published: (2025)
Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
by: Nagori, Aditya, et al.
Published: (2025)
by: Nagori, Aditya, et al.
Published: (2025)
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
by: Hardy, Michael
Published: (2024)
by: Hardy, Michael
Published: (2024)
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
The Use of a Large Language Model for Cyberbullying Detection
by: Ogunleye, Bayode, et al.
Published: (2024)
by: Ogunleye, Bayode, et al.
Published: (2024)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning
by: Jin, Congyun, et al.
Published: (2024)
by: Jin, Congyun, et al.
Published: (2024)
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
by: Oleson, Jon
Published: (2024)
by: Oleson, Jon
Published: (2024)
Improving LLM Leaderboards with Psychometrical Methodology
by: Federiakin, Denis
Published: (2025)
by: Federiakin, Denis
Published: (2025)
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
by: Davoudi, Seyed Pouyan Mousavi, et al.
Published: (2025)
by: Davoudi, Seyed Pouyan Mousavi, et al.
Published: (2025)
A Bayesian Approach to Harnessing the Power of LLMs in Authorship Attribution
by: Hu, Zhengmian, et al.
Published: (2024)
by: Hu, Zhengmian, et al.
Published: (2024)
Language-Dependent Political Bias in AI: A Study of ChatGPT and Gemini
by: Yuksel, Dogus, et al.
Published: (2025)
by: Yuksel, Dogus, et al.
Published: (2025)
Dissecting Physics Reasoning in Small Language Models: A Multi-Dimensional Analysis from an Educational Perspective
by: Scaria, Nicy, et al.
Published: (2025)
by: Scaria, Nicy, et al.
Published: (2025)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
LLM4ED: Large Language Models for Automatic Equation Discovery
by: Du, Mengge, et al.
Published: (2024)
by: Du, Mengge, et al.
Published: (2024)
United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
by: von der Heyde, Leah, et al.
Published: (2024)
by: von der Heyde, Leah, et al.
Published: (2024)
Performance Evaluation of Large Language Models in Statistical Programming
by: Song, Xinyi, et al.
Published: (2025)
by: Song, Xinyi, et al.
Published: (2025)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
by: Zhou, Cai, et al.
Published: (2026)
by: Zhou, Cai, et al.
Published: (2026)
Beyond the Hype: Embeddings vs. Prompting for Multiclass Classification Tasks
by: Kokkodis, Marios, et al.
Published: (2025)
by: Kokkodis, Marios, et al.
Published: (2025)
ICE-ID: A Novel Historical Census Dataset for Longitudinal Identity Resolution
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Unified Representation of Genomic and Biomedical Concepts through Multi-Task, Multi-Source Contrastive Learning
by: Yuan, Hongyi, et al.
Published: (2024)
by: Yuan, Hongyi, et al.
Published: (2024)
A Rational Analysis of the Speech-to-Song Illusion
by: Marjieh, Raja, et al.
Published: (2024)
by: Marjieh, Raja, et al.
Published: (2024)
Crowdsourced Adaptive Surveys
by: Velez, Yamil
Published: (2024)
by: Velez, Yamil
Published: (2024)
Addressing Longstanding Challenges in Cognitive Science with Language Models
by: Wulff, Dirk U., et al.
Published: (2025)
by: Wulff, Dirk U., et al.
Published: (2025)
StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis
by: Song, Xinyi, et al.
Published: (2025)
by: Song, Xinyi, et al.
Published: (2025)
Similar Items
-
Evaluating NLP Embedding Models for Handling Science-Specific Symbolic Expressions in Student Texts
by: Bleckmann, Tom, et al.
Published: (2025) -
Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment
by: Tschisgale, Paul, et al.
Published: (2025) -
Developing and Evaluating a Large Language Model-Based Automated Feedback System Grounded in Evidence-Centered Design for Supporting Physics Problem Solving
by: Maus, Holger, et al.
Published: (2025) -
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025) -
Metacognitive Myopia in Large Language Models
by: Scholten, Florian, et al.
Published: (2024)