Evaluating Large Language Models on the 2026 Korean CSAT Mathematics Exam: Measuring Mathematical Ability in a Zero-Data-Leakage Setting
Fuente:
arXiv
Salvato in:
| Autori principali: | Pyeon, Goun, Heo, Inbum, Jung, Jeesu, Hwang, Taewook, Namgoong, Hyuk, Seo, Hyein, Han, Yerim, Kim, Eunbin, Kang, Hyeonseok, Jung, Sangkeun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
di: Heo, Inbum, et al.
Pubblicazione: (2026)
di: Heo, Inbum, et al.
Pubblicazione: (2026)
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
di: Heo, Inbum, et al.
Pubblicazione: (2025)
di: Heo, Inbum, et al.
Pubblicazione: (2025)
What Drives Paper Acceptance? A Process-Centric Analysis of Modern Peer Review
di: Jung, Sangkeun, et al.
Pubblicazione: (2025)
di: Jung, Sangkeun, et al.
Pubblicazione: (2025)
Exploring Domain Robust Lightweight Reward Models based on Router Mechanism
di: Namgoong, Hyuk, et al.
Pubblicazione: (2024)
di: Namgoong, Hyuk, et al.
Pubblicazione: (2024)
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
di: Hwang, Taewook, et al.
Pubblicazione: (2024)
di: Hwang, Taewook, et al.
Pubblicazione: (2024)
Guidance-Based Prompt Data Augmentation in Specialized Domains for Named Entity Recognition
di: Kang, Hyeonseok, et al.
Pubblicazione: (2024)
di: Kang, Hyeonseok, et al.
Pubblicazione: (2024)
Reasoning Steps as Curriculum: Using Depth of Thought as a Difficulty Signal for Tuning LLMs
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
FEAT: A Preference Feedback Dataset through a Cost-Effective Auto-Generation and Labeling Framework for English AI Tutoring
di: Seo, Hyein, et al.
Pubblicazione: (2025)
di: Seo, Hyein, et al.
Pubblicazione: (2025)
A Pilot Study on the Short‐Term Effects of an Electric Knee–Ankle–Foot Orthosis on Gait Performance and Physiological Cost Index in Patients With Hemiplegia: Influence of Initial Balance Ability Assessed by the Berg Balance Scale
di: Hyuk-Jae Choi, et al.
Pubblicazione: (2026)
di: Hyuk-Jae Choi, et al.
Pubblicazione: (2026)
Does Incomplete Syntax Influence Korean Language Model? Focusing on Word Order and Case Markers
di: Kim, Jong Myoung, et al.
Pubblicazione: (2024)
di: Kim, Jong Myoung, et al.
Pubblicazione: (2024)
Phasor Memory Networks: Stable Backpropagation Through Time for Scalable Explicit Memory
di: Goo, Sungwoo, et al.
Pubblicazione: (2026)
di: Goo, Sungwoo, et al.
Pubblicazione: (2026)
SC-Lane: Slope-aware and Consistent Road Height Estimation Framework for 3D Lane Detection
di: Park, Chaesong, et al.
Pubblicazione: (2025)
di: Park, Chaesong, et al.
Pubblicazione: (2025)
On the robustness of ChatGPT in teaching Korean Mathematics
di: Nguyen, Phuong-Nam, et al.
Pubblicazione: (2025)
di: Nguyen, Phuong-Nam, et al.
Pubblicazione: (2025)
Won: Establishing Best Practices for Korean Financial NLP
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
Unraveling the Complexities of Hepatitis B Virus Control: A Mathematical Modeling Approach
di: Malik Muhammad Ibrahim, et al.
Pubblicazione: (2025)
di: Malik Muhammad Ibrahim, et al.
Pubblicazione: (2025)
MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
di: Hyeon, Sieun, et al.
Pubblicazione: (2024)
di: Hyeon, Sieun, et al.
Pubblicazione: (2024)
Pattern of Metastasis as a Risk Factor for Brain Metastasis in Patients With Breast Cancer: A Time‐Dependent Cox Regression
di: Eunbin Park, et al.
Pubblicazione: (2026)
di: Eunbin Park, et al.
Pubblicazione: (2026)
Technology Innovations and the Development of Distance Education: Korean Experience.
di: Jung, Insung
Pubblicazione: (2000)
di: Jung, Insung
Pubblicazione: (2000)
HeightLane: BEV Heightmap guided 3D Lane Detection
di: Park, Chaesong, et al.
Pubblicazione: (2024)
di: Park, Chaesong, et al.
Pubblicazione: (2024)
Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models
di: Lee, Jung Hyun, et al.
Pubblicazione: (2024)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2024)
Towards Robust Mathematical Reasoning
di: Luong, Thang, et al.
Pubblicazione: (2025)
di: Luong, Thang, et al.
Pubblicazione: (2025)
AlephZero and Mathematical Experience
di: DeDeo, Simon
Pubblicazione: (2023)
di: DeDeo, Simon
Pubblicazione: (2023)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
di: Hwang, Hyeonbin, et al.
Pubblicazione: (2024)
di: Hwang, Hyeonbin, et al.
Pubblicazione: (2024)
Health Benefits of Family Visits for Older Korean Women Living Alone
di: Soo‐Ji Hwang, et al.
Pubblicazione: (2026)
di: Soo‐Ji Hwang, et al.
Pubblicazione: (2026)
Review of the genus Elasmostethus (Hemiptera: Heteroptera: Acanthosomatidae) from the Korean Peninsula
di: Jung, Sunghoon
Pubblicazione: (2017)
di: Jung, Sunghoon
Pubblicazione: (2017)
LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model
di: Jang, Youngjoon, et al.
Pubblicazione: (2026)
di: Jang, Youngjoon, et al.
Pubblicazione: (2026)
MathReader : Text-to-Speech for Mathematical Documents
di: Hyeon, Sieun, et al.
Pubblicazione: (2025)
di: Hyeon, Sieun, et al.
Pubblicazione: (2025)
Pride, Not Prejudice
di: Chung, Eunbin
Pubblicazione: (2022)
di: Chung, Eunbin
Pubblicazione: (2022)
Pride, Not Prejudice
di: Chung, Eunbin
Pubblicazione: (2022)
di: Chung, Eunbin
Pubblicazione: (2022)
Evaluating the Reasoning Abilities of LLMs on Underrepresented Mathematics Competition Problems
di: Golladay, Samuel, et al.
Pubblicazione: (2025)
di: Golladay, Samuel, et al.
Pubblicazione: (2025)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
di: Kuang, Jiayi, et al.
Pubblicazione: (2025)
di: Kuang, Jiayi, et al.
Pubblicazione: (2025)
AI-assisted Automated Short Answer Grading of Handwritten University Level Mathematics Exams
di: Liu, Tianyi, et al.
Pubblicazione: (2024)
di: Liu, Tianyi, et al.
Pubblicazione: (2024)
CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam
di: Khrulev, Ruslan
Pubblicazione: (2025)
di: Khrulev, Ruslan
Pubblicazione: (2025)
V-Math: An Agentic Approach to the Vietnamese National High School Graduation Mathematics Exams
di: Nguyen, Duong Q., et al.
Pubblicazione: (2025)
di: Nguyen, Duong Q., et al.
Pubblicazione: (2025)
MathDoc: Benchmarking Structured Extraction and Active Refusal on Noisy Mathematics Exam Papers
di: Zhou, Chenyue, et al.
Pubblicazione: (2026)
di: Zhou, Chenyue, et al.
Pubblicazione: (2026)
PAD: Towards Efficient Data Generation for Transfer Learning Using Phrase Alignment
di: Kim, Jong Myoung, et al.
Pubblicazione: (2025)
di: Kim, Jong Myoung, et al.
Pubblicazione: (2025)
LANGALIGN: Enhancing Non-English Language Models via Cross-Lingual Embedding Alignment
di: Kim, Jong Myoung, et al.
Pubblicazione: (2025)
di: Kim, Jong Myoung, et al.
Pubblicazione: (2025)
MathBridge: A Large Corpus Dataset for Translating Spoken Mathematical Expressions into $LaTeX$ Formulas for Improved Readability
di: Jung, Kyudan, et al.
Pubblicazione: (2024)
di: Jung, Kyudan, et al.
Pubblicazione: (2024)
Korean Christian Missionaries in High‐Risk Countries: Interaction with International Religious Networks and Domestic Response
di: Jihye Jung
Pubblicazione: (2024)
di: Jihye Jung
Pubblicazione: (2024)
Documenti analoghi
-
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
di: Heo, Inbum, et al.
Pubblicazione: (2026) -
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
di: Heo, Inbum, et al.
Pubblicazione: (2025) -
What Drives Paper Acceptance? A Process-Centric Analysis of Modern Peer Review
di: Jung, Sangkeun, et al.
Pubblicazione: (2025) -
Exploring Domain Robust Lightweight Reward Models based on Router Mechanism
di: Namgoong, Hyuk, et al.
Pubblicazione: (2024) -
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
di: Hwang, Taewook, et al.
Pubblicazione: (2024)