Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Jiaxing, Huang, Weiquan, Wu, Jiang, Gu, Chenya, Li, Wei, Zhang, Songyang, Yan, Hang, He, Conghui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
von: You, Wangjie, et al.
Veröffentlicht: (2025)
von: You, Wangjie, et al.
Veröffentlicht: (2025)
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
von: Wang, Zhenglin, et al.
Veröffentlicht: (2025)
von: Wang, Zhenglin, et al.
Veröffentlicht: (2025)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
von: Xu, Liang, et al.
Veröffentlicht: (2024)
von: Xu, Liang, et al.
Veröffentlicht: (2024)
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
von: Zhang, Ming, et al.
Veröffentlicht: (2026)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge Reasoning
von: Cui, Shaobo, et al.
Veröffentlicht: (2024)
von: Cui, Shaobo, et al.
Veröffentlicht: (2024)
Flames: Benchmarking Value Alignment of LLMs in Chinese
von: Huang, Kexin, et al.
Veröffentlicht: (2023)
von: Huang, Kexin, et al.
Veröffentlicht: (2023)
Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2024)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
von: Kim, Jisu, et al.
Veröffentlicht: (2025)
von: Kim, Jisu, et al.
Veröffentlicht: (2025)
HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning
von: Junias, Obed, et al.
Veröffentlicht: (2026)
von: Junias, Obed, et al.
Veröffentlicht: (2026)
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
von: Zhan, Weidong, et al.
Veröffentlicht: (2025)
von: Zhan, Weidong, et al.
Veröffentlicht: (2025)
Thai Winograd Schemas: A Benchmark for Thai Commonsense Reasoning
von: Artkaew, Phakphum
Veröffentlicht: (2024)
von: Artkaew, Phakphum
Veröffentlicht: (2024)
CMDAR: A Chinese Multi-scene Dynamic Audio Reasoning Benchmark with Diverse Challenges
von: Li, Hui, et al.
Veröffentlicht: (2025)
von: Li, Hui, et al.
Veröffentlicht: (2025)
Multi-Physics: A Comprehensive Benchmark for Multimodal LLMs Reasoning on Chinese Multi-Subject Physics Problems
von: Luo, Zhongze, et al.
Veröffentlicht: (2025)
von: Luo, Zhongze, et al.
Veröffentlicht: (2025)
Commonsense Reasoning in Arab Culture
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
von: Kong, Shufeng, et al.
Veröffentlicht: (2025)
von: Kong, Shufeng, et al.
Veröffentlicht: (2025)
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
von: Wu, Mingqi, et al.
Veröffentlicht: (2025)
von: Wu, Mingqi, et al.
Veröffentlicht: (2025)
From Memorization to Reasoning in the Spectrum of Loss Curvature
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
Reason to Rote: Rethinking Memorization in Reasoning
von: Du, Yupei, et al.
Veröffentlicht: (2025)
von: Du, Yupei, et al.
Veröffentlicht: (2025)
GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
von: Zhu, Shaojie, et al.
Veröffentlicht: (2023)
von: Zhu, Shaojie, et al.
Veröffentlicht: (2023)
The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation
von: Diallo, Aissatou, et al.
Veröffentlicht: (2025)
von: Diallo, Aissatou, et al.
Veröffentlicht: (2025)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
von: Li, Aochong Oliver, et al.
Veröffentlicht: (2025)
Quantifying In-Context Reasoning Effects and Memorization Effects in LLMs
von: Lou, Siyu, et al.
Veröffentlicht: (2024)
von: Lou, Siyu, et al.
Veröffentlicht: (2024)
On Memorization of Large Language Models in Logical Reasoning
von: Xie, Chulin, et al.
Veröffentlicht: (2024)
von: Xie, Chulin, et al.
Veröffentlicht: (2024)
Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
CSSBench: Evaluating the Safety of Lightweight LLMs against Chinese-Specific Adversarial Patterns
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2026)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2026)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
Evaluating LLMs on Chinese Idiom Translation
von: Yang, Cai, et al.
Veröffentlicht: (2025)
von: Yang, Cai, et al.
Veröffentlicht: (2025)
Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning
von: Qiu, Ziqi, et al.
Veröffentlicht: (2024)
von: Qiu, Ziqi, et al.
Veröffentlicht: (2024)
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation
von: Hong, Mengze, et al.
Veröffentlicht: (2025)
von: Hong, Mengze, et al.
Veröffentlicht: (2025)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
von: Li, Wei, et al.
Veröffentlicht: (2024) -
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
von: You, Wangjie, et al.
Veröffentlicht: (2025) -
Benchmarking Temporal Reasoning and Alignment Across Chinese Dynasties
von: Wang, Zhenglin, et al.
Veröffentlicht: (2025) -
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
von: Xu, Liang, et al.
Veröffentlicht: (2024) -
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
von: Zhang, Ming, et al.
Veröffentlicht: (2026)