Evaluating Large Language Models on the 2026 Korean CSAT Mathematics Exam: Measuring Mathematical Ability in a Zero-Data-Leakage Setting
Fuente:
arXiv
Saved in:
| Main Authors: | Pyeon, Goun, Heo, Inbum, Jung, Jeesu, Hwang, Taewook, Namgoong, Hyuk, Seo, Hyein, Han, Yerim, Kim, Eunbin, Kang, Hyeonseok, Jung, Sangkeun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
by: Heo, Inbum, et al.
Published: (2026)
by: Heo, Inbum, et al.
Published: (2026)
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
by: Heo, Inbum, et al.
Published: (2025)
by: Heo, Inbum, et al.
Published: (2025)
What Drives Paper Acceptance? A Process-Centric Analysis of Modern Peer Review
by: Jung, Sangkeun, et al.
Published: (2025)
by: Jung, Sangkeun, et al.
Published: (2025)
Exploring Domain Robust Lightweight Reward Models based on Router Mechanism
by: Namgoong, Hyuk, et al.
Published: (2024)
by: Namgoong, Hyuk, et al.
Published: (2024)
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
by: Hwang, Taewook, et al.
Published: (2024)
by: Hwang, Taewook, et al.
Published: (2024)
Guidance-Based Prompt Data Augmentation in Specialized Domains for Named Entity Recognition
by: Kang, Hyeonseok, et al.
Published: (2024)
by: Kang, Hyeonseok, et al.
Published: (2024)
Reasoning Steps as Curriculum: Using Depth of Thought as a Difficulty Signal for Tuning LLMs
by: Jung, Jeesu, et al.
Published: (2025)
by: Jung, Jeesu, et al.
Published: (2025)
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
by: Jung, Jeesu, et al.
Published: (2025)
by: Jung, Jeesu, et al.
Published: (2025)
FEAT: A Preference Feedback Dataset through a Cost-Effective Auto-Generation and Labeling Framework for English AI Tutoring
by: Seo, Hyein, et al.
Published: (2025)
by: Seo, Hyein, et al.
Published: (2025)
A Pilot Study on the Short‐Term Effects of an Electric Knee–Ankle–Foot Orthosis on Gait Performance and Physiological Cost Index in Patients With Hemiplegia: Influence of Initial Balance Ability Assessed by the Berg Balance Scale
by: Hyuk-Jae Choi, et al.
Published: (2026)
by: Hyuk-Jae Choi, et al.
Published: (2026)
Does Incomplete Syntax Influence Korean Language Model? Focusing on Word Order and Case Markers
by: Kim, Jong Myoung, et al.
Published: (2024)
by: Kim, Jong Myoung, et al.
Published: (2024)
Phasor Memory Networks: Stable Backpropagation Through Time for Scalable Explicit Memory
by: Goo, Sungwoo, et al.
Published: (2026)
by: Goo, Sungwoo, et al.
Published: (2026)
SC-Lane: Slope-aware and Consistent Road Height Estimation Framework for 3D Lane Detection
by: Park, Chaesong, et al.
Published: (2025)
by: Park, Chaesong, et al.
Published: (2025)
On the robustness of ChatGPT in teaching Korean Mathematics
by: Nguyen, Phuong-Nam, et al.
Published: (2025)
by: Nguyen, Phuong-Nam, et al.
Published: (2025)
Won: Establishing Best Practices for Korean Financial NLP
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Unraveling the Complexities of Hepatitis B Virus Control: A Mathematical Modeling Approach
by: Malik Muhammad Ibrahim, et al.
Published: (2025)
by: Malik Muhammad Ibrahim, et al.
Published: (2025)
MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
by: Hyeon, Sieun, et al.
Published: (2024)
by: Hyeon, Sieun, et al.
Published: (2024)
Pattern of Metastasis as a Risk Factor for Brain Metastasis in Patients With Breast Cancer: A Time‐Dependent Cox Regression
by: Eunbin Park, et al.
Published: (2026)
by: Eunbin Park, et al.
Published: (2026)
Technology Innovations and the Development of Distance Education: Korean Experience.
by: Jung, Insung
Published: (2000)
by: Jung, Insung
Published: (2000)
HeightLane: BEV Heightmap guided 3D Lane Detection
by: Park, Chaesong, et al.
Published: (2024)
by: Park, Chaesong, et al.
Published: (2024)
Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models
by: Lee, Jung Hyun, et al.
Published: (2024)
by: Lee, Jung Hyun, et al.
Published: (2024)
Towards Robust Mathematical Reasoning
by: Luong, Thang, et al.
Published: (2025)
by: Luong, Thang, et al.
Published: (2025)
AlephZero and Mathematical Experience
by: DeDeo, Simon
Published: (2023)
by: DeDeo, Simon
Published: (2023)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
by: Hwang, Hyeonbin, et al.
Published: (2024)
by: Hwang, Hyeonbin, et al.
Published: (2024)
Health Benefits of Family Visits for Older Korean Women Living Alone
by: Soo‐Ji Hwang, et al.
Published: (2026)
by: Soo‐Ji Hwang, et al.
Published: (2026)
Review of the genus Elasmostethus (Hemiptera: Heteroptera: Acanthosomatidae) from the Korean Peninsula
by: Jung, Sunghoon
Published: (2017)
by: Jung, Sunghoon
Published: (2017)
LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model
by: Jang, Youngjoon, et al.
Published: (2026)
by: Jang, Youngjoon, et al.
Published: (2026)
MathReader : Text-to-Speech for Mathematical Documents
by: Hyeon, Sieun, et al.
Published: (2025)
by: Hyeon, Sieun, et al.
Published: (2025)
Pride, Not Prejudice
by: Chung, Eunbin
Published: (2022)
by: Chung, Eunbin
Published: (2022)
Pride, Not Prejudice
by: Chung, Eunbin
Published: (2022)
by: Chung, Eunbin
Published: (2022)
Evaluating the Reasoning Abilities of LLMs on Underrepresented Mathematics Competition Problems
by: Golladay, Samuel, et al.
Published: (2025)
by: Golladay, Samuel, et al.
Published: (2025)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
by: Kuang, Jiayi, et al.
Published: (2025)
by: Kuang, Jiayi, et al.
Published: (2025)
AI-assisted Automated Short Answer Grading of Handwritten University Level Mathematics Exams
by: Liu, Tianyi, et al.
Published: (2024)
by: Liu, Tianyi, et al.
Published: (2024)
CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam
by: Khrulev, Ruslan
Published: (2025)
by: Khrulev, Ruslan
Published: (2025)
V-Math: An Agentic Approach to the Vietnamese National High School Graduation Mathematics Exams
by: Nguyen, Duong Q., et al.
Published: (2025)
by: Nguyen, Duong Q., et al.
Published: (2025)
MathDoc: Benchmarking Structured Extraction and Active Refusal on Noisy Mathematics Exam Papers
by: Zhou, Chenyue, et al.
Published: (2026)
by: Zhou, Chenyue, et al.
Published: (2026)
PAD: Towards Efficient Data Generation for Transfer Learning Using Phrase Alignment
by: Kim, Jong Myoung, et al.
Published: (2025)
by: Kim, Jong Myoung, et al.
Published: (2025)
LANGALIGN: Enhancing Non-English Language Models via Cross-Lingual Embedding Alignment
by: Kim, Jong Myoung, et al.
Published: (2025)
by: Kim, Jong Myoung, et al.
Published: (2025)
MathBridge: A Large Corpus Dataset for Translating Spoken Mathematical Expressions into $LaTeX$ Formulas for Improved Readability
by: Jung, Kyudan, et al.
Published: (2024)
by: Jung, Kyudan, et al.
Published: (2024)
Korean Christian Missionaries in High‐Risk Countries: Interaction with International Religious Networks and Domestic Response
by: Jihye Jung
Published: (2024)
by: Jihye Jung
Published: (2024)
Similar Items
-
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
by: Heo, Inbum, et al.
Published: (2026) -
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
by: Heo, Inbum, et al.
Published: (2025) -
What Drives Paper Acceptance? A Process-Centric Analysis of Modern Peer Review
by: Jung, Sangkeun, et al.
Published: (2025) -
Exploring Domain Robust Lightweight Reward Models based on Router Mechanism
by: Namgoong, Hyuk, et al.
Published: (2024) -
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
by: Hwang, Taewook, et al.
Published: (2024)