GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912633034113024 |
|---|---|
| author | Zhang, Yue Zhang, Jiaxin Ren, Qiuyu Saffat, Tahsin Liu, Xiaoxuan Yang, Zitong Zhu, Banghua Ma, Yi |
| author_facet | Zhang, Yue Zhang, Jiaxin Ren, Qiuyu Saffat, Tahsin Liu, Xiaoxuan Yang, Zitong Zhu, Banghua Ma, Yi |
| contents | We introduce \textbf{GAUSS} (\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics), a benchmark that evaluates LLMs' mathematical abilities across twelve core skill dimensions, grouped into three domains: knowledge and understanding, problem solving and communication, and meta-skills and creativity. By categorizing problems according to cognitive skills and designing tasks that isolate specific abilities, GAUSS constructs comprehensive, fine-grained, and interpretable profiles of models' mathematical abilities. These profiles faithfully represent their underlying mathematical intelligence. To exemplify how to use the \textsc{GAUSS} benchmark, we have derived the skill profile of \textsc{GPT-5-thinking}, revealing its strengths and weaknesses as well as its differences relative to \textsc{o4-mini-high}, thereby underscoring the value of multidimensional, skill-based evaluation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_18122 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models Zhang, Yue Zhang, Jiaxin Ren, Qiuyu Saffat, Tahsin Liu, Xiaoxuan Yang, Zitong Zhu, Banghua Ma, Yi Artificial Intelligence Computation and Language We introduce \textbf{GAUSS} (\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics), a benchmark that evaluates LLMs' mathematical abilities across twelve core skill dimensions, grouped into three domains: knowledge and understanding, problem solving and communication, and meta-skills and creativity. By categorizing problems according to cognitive skills and designing tasks that isolate specific abilities, GAUSS constructs comprehensive, fine-grained, and interpretable profiles of models' mathematical abilities. These profiles faithfully represent their underlying mathematical intelligence. To exemplify how to use the \textsc{GAUSS} benchmark, we have derived the skill profile of \textsc{GPT-5-thinking}, revealing its strengths and weaknesses as well as its differences relative to \textsc{o4-mini-high}, thereby underscoring the value of multidimensional, skill-based evaluation. |
| title | GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2509.18122 |