SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
Fuente:
arXiv
Salvato in:
| Autori principali: | Dinh, Tu Anh, Mullov, Carlos, Bärmann, Leonard, Li, Zhaolin, Liu, Danni, Reiß, Simon, Lee, Jueun, Lerzer, Nathan, Ternava, Fabian, Gao, Jianfeng, Röddiger, Tobias, Waibel, Alexander, Asfour, Tamim, Beigl, Michael, Stiefelhagen, Rainer, Dachsbacher, Carsten, Böhm, Klemens, Niehues, Jan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning to Forget -- Hierarchical Episodic Memory for Lifelong Robot Deployment
di: Bärmann, Leonard, et al.
Pubblicazione: (2026)
di: Bärmann, Leonard, et al.
Pubblicazione: (2026)
Incremental Learning of Humanoid Robot Behavior from Natural Interaction and Large Language Models
di: Bärmann, Leonard, et al.
Pubblicazione: (2023)
di: Bärmann, Leonard, et al.
Pubblicazione: (2023)
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
di: Li, Zhaolin, et al.
Pubblicazione: (2025)
di: Li, Zhaolin, et al.
Pubblicazione: (2025)
Haptic Biofeedback for Wakeful Rest: Does Stimulation Location Make a Difference?
di: Lee, Jueun, et al.
Pubblicazione: (2025)
di: Lee, Jueun, et al.
Pubblicazione: (2025)
Exploring LLMs for Scientific Information Extraction Using The SciEx Framework
di: Li, Sha, et al.
Pubblicazione: (2025)
di: Li, Sha, et al.
Pubblicazione: (2025)
Feasibility of In-Ear Single-Channel ExG for Wearable Sleep Monitoring in Real-World Settings
di: Lepold, Philipp, et al.
Pubblicazione: (2025)
di: Lepold, Philipp, et al.
Pubblicazione: (2025)
Episodic Memory Verbalization using Hierarchical Representations of Life-Long Robot Experience
di: Bärmann, Leonard, et al.
Pubblicazione: (2024)
di: Bärmann, Leonard, et al.
Pubblicazione: (2024)
Closed-Loop Rhythmic Haptic Biofeedback via Smartwatch for Relaxation and Sleep Onset
di: Lee, Jueun, et al.
Pubblicazione: (2025)
di: Lee, Jueun, et al.
Pubblicazione: (2025)
HARNode: A Time-Synchronised, Open-Source, Multi-Device, Wearable System for Ad Hoc Field Studies
di: Lepold, Philipp, et al.
Pubblicazione: (2025)
di: Lepold, Philipp, et al.
Pubblicazione: (2025)
Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen Languages
di: Mullov, Carlos, et al.
Pubblicazione: (2024)
di: Mullov, Carlos, et al.
Pubblicazione: (2024)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
di: Huber, Christian, et al.
Pubblicazione: (2023)
di: Huber, Christian, et al.
Pubblicazione: (2023)
OpenEarable ExG: Open-Source Hardware for Ear-Based Biopotential Sensing Applications
di: Lepold, Philipp, et al.
Pubblicazione: (2024)
di: Lepold, Philipp, et al.
Pubblicazione: (2024)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
di: Koneru, Sai, et al.
Pubblicazione: (2024)
di: Koneru, Sai, et al.
Pubblicazione: (2024)
Visual Imitation Learning of Task-Oriented Object Grasping and Rearrangement
di: Cai, Yichen, et al.
Pubblicazione: (2024)
di: Cai, Yichen, et al.
Pubblicazione: (2024)
EarCapAuth: Biometric Method for Earables Using Capacitive Sensing Eartips
di: Hanser, Richard, et al.
Pubblicazione: (2024)
di: Hanser, Richard, et al.
Pubblicazione: (2024)
MicroNAS: Memory and Latency Constrained Hardware-Aware Neural Architecture Search for Time Series Classification on Microcontrollers
di: King, Tobias, et al.
Pubblicazione: (2023)
di: King, Tobias, et al.
Pubblicazione: (2023)
pcbGPT: Automatic PCB Schematic Synthesis from Natural Language Requirements
di: King, Tobias, et al.
Pubblicazione: (2026)
di: King, Tobias, et al.
Pubblicazione: (2026)
Learning Symbolic and Subsymbolic Temporal Task Constraints from Bimanual Human Demonstrations
di: Dreher, Christian, et al.
Pubblicazione: (2024)
di: Dreher, Christian, et al.
Pubblicazione: (2024)
KITchen: A Real-World Benchmark and Dataset for 6D Object Pose Estimation in Kitchen Environments
di: Younes, Abdelrahman, et al.
Pubblicazione: (2024)
di: Younes, Abdelrahman, et al.
Pubblicazione: (2024)
A Flexible Architecture for Privacy-Aware Trust Management
di: Klemens Böhm
Pubblicazione: (2010)
di: Klemens Böhm
Pubblicazione: (2010)
Feature Toggle Dynamics in Large-Scale Systems: Prevalence, Growth, Lifespan, and Benchmarking
di: Tërnava, Xhevahire
Pubblicazione: (2026)
di: Tërnava, Xhevahire
Pubblicazione: (2026)
In-context Language Learning for Endangered Languages in Speech Recognition
di: Li, Zhaolin, et al.
Pubblicazione: (2025)
di: Li, Zhaolin, et al.
Pubblicazione: (2025)
Multimodal In-context Learning for ASR of Low-resource Languages
di: Li, Zhaolin, et al.
Pubblicazione: (2026)
di: Li, Zhaolin, et al.
Pubblicazione: (2026)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
di: Zink, Oswald, et al.
Pubblicazione: (2024)
di: Zink, Oswald, et al.
Pubblicazione: (2024)
Conditions for Catastrophic Forgetting in Multilingual Translation
di: Liu, Danni, et al.
Pubblicazione: (2025)
di: Liu, Danni, et al.
Pubblicazione: (2025)
Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs
di: Liu, Danni, et al.
Pubblicazione: (2025)
di: Liu, Danni, et al.
Pubblicazione: (2025)
How Transferable are Attribute Controllers on Pretrained Multilingual Translation Models?
di: Liu, Danni, et al.
Pubblicazione: (2023)
di: Liu, Danni, et al.
Pubblicazione: (2023)
Overview of AI Grading of Physics Olympiad Exams
di: McGinness, Lachlan
Pubblicazione: (2025)
di: McGinness, Lachlan
Pubblicazione: (2025)
Bi-KVIL: Keypoints-based Visual Imitation Learning of Bimanual Manipulation Tasks
di: Gao, Jianfeng, et al.
Pubblicazione: (2024)
di: Gao, Jianfeng, et al.
Pubblicazione: (2024)
WHAR Datasets: An Open Source Library for Wearable Human Activity Recognition
di: Burzer, Maximilian, et al.
Pubblicazione: (2025)
di: Burzer, Maximilian, et al.
Pubblicazione: (2025)
UltrasonicSpheres: Localized, Multi-Channel Sound Spheres Using Off-the-Shelf Speakers and Earables
di: Küttner, Michael, et al.
Pubblicazione: (2025)
di: Küttner, Michael, et al.
Pubblicazione: (2025)
EarResp-ANS : Audio-Based On-Device Respiration Rate Estimation on Earphones with Adaptive Noise Suppression
di: Küttner, Michael, et al.
Pubblicazione: (2026)
di: Küttner, Michael, et al.
Pubblicazione: (2026)
Unraveling the Single Tangent Space Fallacy: An Analysis and Clarification for Applying Riemannian Geometry in Robot Learning
di: Jaquier, Noémie, et al.
Pubblicazione: (2023)
di: Jaquier, Noémie, et al.
Pubblicazione: (2023)
Adaptive Domain Modeling with Language Models: A Multi-Agent Approach to Task Planning
di: Babu, Harisankar, et al.
Pubblicazione: (2025)
di: Babu, Harisankar, et al.
Pubblicazione: (2025)
Towards continually learning new languages
di: Pham, Ngoc-Quan, et al.
Pubblicazione: (2022)
di: Pham, Ngoc-Quan, et al.
Pubblicazione: (2022)
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
Grading Handwritten Engineering Exams with Multimodal Large Language Models
di: Perš, Janez, et al.
Pubblicazione: (2026)
di: Perš, Janez, et al.
Pubblicazione: (2026)
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
di: Caraeni, Adriana, et al.
Pubblicazione: (2024)
di: Caraeni, Adriana, et al.
Pubblicazione: (2024)
Evaluating Privacy Measures for Load Hiding
di: Arzamasov, Vadim, et al.
Pubblicazione: (2024)
di: Arzamasov, Vadim, et al.
Pubblicazione: (2024)
Heatables: Effects of Infrared-LED-Induced Ear Heating on Thermal Perception, Comfort, and Cognitive Performance
di: Zitz, Valeria, et al.
Pubblicazione: (2025)
di: Zitz, Valeria, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Learning to Forget -- Hierarchical Episodic Memory for Lifelong Robot Deployment
di: Bärmann, Leonard, et al.
Pubblicazione: (2026) -
Incremental Learning of Humanoid Robot Behavior from Natural Interaction and Large Language Models
di: Bärmann, Leonard, et al.
Pubblicazione: (2023) -
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
di: Li, Zhaolin, et al.
Pubblicazione: (2025) -
Haptic Biofeedback for Wakeful Rest: Does Stimulation Location Make a Difference?
di: Lee, Jueun, et al.
Pubblicazione: (2025) -
Exploring LLMs for Scientific Information Extraction Using The SciEx Framework
di: Li, Sha, et al.
Pubblicazione: (2025)