GLoRE: Evaluating Logical Reasoning of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908327049428992 |
|---|---|
| author | liu, Hanmeng Teng, Zhiyang Ning, Ruoxi Ding, Yiran Li, Xiulai Liu, Xiaozhang Zhang, Yue |
| author_facet | liu, Hanmeng Teng, Zhiyang Ning, Ruoxi Ding, Yiran Li, Xiulai Liu, Xiaozhang Zhang, Yue |
| contents | Large language models (LLMs) have shown significant general language understanding abilities. However, there has been a scarcity of attempts to assess the logical reasoning capacities of these LLMs, an essential facet of natural language understanding. To encourage further investigation in this area, we introduce GLoRE, a General Logical Reasoning Evaluation platform that not only consolidates diverse datasets but also standardizes them into a unified format suitable for evaluating large language models across zero-shot and few-shot scenarios. Our experimental results show that compared to the performance of humans and supervised fine-tuning models, the logical reasoning capabilities of large reasoning models, such as OpenAI's o1 mini, DeepSeek R1 and QwQ-32B, have seen remarkable improvements, with QwQ-32B achieving the highest benchmark performance to date. GLoRE is designed as a living project that continuously integrates new datasets and models, facilitating robust and comparative assessments of model performance in both commercial and Huggingface communities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_09107 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | GLoRE: Evaluating Logical Reasoning of Large Language Models liu, Hanmeng Teng, Zhiyang Ning, Ruoxi Ding, Yiran Li, Xiulai Liu, Xiaozhang Zhang, Yue Computation and Language Artificial Intelligence Large language models (LLMs) have shown significant general language understanding abilities. However, there has been a scarcity of attempts to assess the logical reasoning capacities of these LLMs, an essential facet of natural language understanding. To encourage further investigation in this area, we introduce GLoRE, a General Logical Reasoning Evaluation platform that not only consolidates diverse datasets but also standardizes them into a unified format suitable for evaluating large language models across zero-shot and few-shot scenarios. Our experimental results show that compared to the performance of humans and supervised fine-tuning models, the logical reasoning capabilities of large reasoning models, such as OpenAI's o1 mini, DeepSeek R1 and QwQ-32B, have seen remarkable improvements, with QwQ-32B achieving the highest benchmark performance to date. GLoRE is designed as a living project that continuously integrates new datasets and models, facilitating robust and comparative assessments of model performance in both commercial and Huggingface communities. |
| title | GLoRE: Evaluating Logical Reasoning of Large Language Models |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2310.09107 |