CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Jiayu, Huang, Zhenya, Dai, Wei, Cheng, Cheng, Wu, Jinze, Sha, Jing, Li, Song, Liu, Qi, Wang, Shijin, Chen, Enhong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
UniCog: Uncovering Cognitive Abilities of LLMs through Latent Mind Space Analysis
di: Liu, Jiayu, et al.
Pubblicazione: (2026)
di: Liu, Jiayu, et al.
Pubblicazione: (2026)
Learning to Solve Geometry Problems via Simulating Human Dual-Reasoning Process
di: Xiao, Tong, et al.
Pubblicazione: (2024)
di: Xiao, Tong, et al.
Pubblicazione: (2024)
Unified Uncertainty Estimation for Cognitive Diagnosis Models
di: Wang, Fei, et al.
Pubblicazione: (2024)
di: Wang, Fei, et al.
Pubblicazione: (2024)
Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
di: Huang, Zhenya, et al.
Pubblicazione: (2025)
di: Huang, Zhenya, et al.
Pubblicazione: (2025)
Bit-mask Robust Contrastive Knowledge Distillation for Unsupervised Semantic Hashing
di: He, Liyang, et al.
Pubblicazione: (2024)
di: He, Liyang, et al.
Pubblicazione: (2024)
End-to-End Graph Flattening Method for Large Language Models
di: Hong, Bin, et al.
Pubblicazione: (2024)
di: Hong, Bin, et al.
Pubblicazione: (2024)
From Objectives to Questions: A Planning-based Framework for Educational Mathematical Question Generation
di: Cheng, Cheng, et al.
Pubblicazione: (2025)
di: Cheng, Cheng, et al.
Pubblicazione: (2025)
Verifying Large Language Models' Reasoning Paths via Correlation Matrix Rank
di: Liu, Jiayu, et al.
Pubblicazione: (2025)
di: Liu, Jiayu, et al.
Pubblicazione: (2025)
What Makes In-context Learning Effective for Mathematical Reasoning: A Theoretical Analysis
di: Liu, Jiayu, et al.
Pubblicazione: (2024)
di: Liu, Jiayu, et al.
Pubblicazione: (2024)
Position: AI Evaluation Should Learn from How We Test Humans
di: Zhuang, Yan, et al.
Pubblicazione: (2023)
di: Zhuang, Yan, et al.
Pubblicazione: (2023)
A Survey of Models for Cognitive Diagnosis: New Developments and Future Directions
di: Wang, Fei, et al.
Pubblicazione: (2024)
di: Wang, Fei, et al.
Pubblicazione: (2024)
CogEvolution: A Human-like Generative Educational Agent to Simulate Student's Cognitive Evolution
di: Zhang, Wei, et al.
Pubblicazione: (2026)
di: Zhang, Wei, et al.
Pubblicazione: (2026)
TestAgent: An Adaptive and Intelligent Expert for Human Assessment
di: Yu, Junhao, et al.
Pubblicazione: (2025)
di: Yu, Junhao, et al.
Pubblicazione: (2025)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
di: Chen, Zui, et al.
Pubblicazione: (2024)
di: Chen, Zui, et al.
Pubblicazione: (2024)
EgoCogNav: Cognition-aware Human Egocentric Navigation
di: Qiu, Zhiwen, et al.
Pubblicazione: (2025)
di: Qiu, Zhiwen, et al.
Pubblicazione: (2025)
CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
di: Liu, Cheng, et al.
Pubblicazione: (2025)
di: Liu, Cheng, et al.
Pubblicazione: (2025)
ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities
di: Dong, Wenhan, et al.
Pubblicazione: (2025)
di: Dong, Wenhan, et al.
Pubblicazione: (2025)
Can LLMs Solve longer Math Word Problems Better?
di: Xu, Xin, et al.
Pubblicazione: (2024)
di: Xu, Xin, et al.
Pubblicazione: (2024)
MMATH: A Multilingual Benchmark for Mathematical Reasoning
di: Luo, Wenyang, et al.
Pubblicazione: (2025)
di: Luo, Wenyang, et al.
Pubblicazione: (2025)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
di: Zhao, Yuze, et al.
Pubblicazione: (2026)
di: Zhao, Yuze, et al.
Pubblicazione: (2026)
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
di: Jing, Zonglei, et al.
Pubblicazione: (2025)
di: Jing, Zonglei, et al.
Pubblicazione: (2025)
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs
di: Cao, Yihan, et al.
Pubblicazione: (2024)
di: Cao, Yihan, et al.
Pubblicazione: (2024)
Survey of Computerized Adaptive Testing: A Machine Learning Perspective
di: Zhuang, Yan, et al.
Pubblicazione: (2024)
di: Zhuang, Yan, et al.
Pubblicazione: (2024)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
di: Li, Xin, et al.
Pubblicazione: (2025)
di: Li, Xin, et al.
Pubblicazione: (2025)
Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
di: Hu, Yijie, et al.
Pubblicazione: (2025)
di: Hu, Yijie, et al.
Pubblicazione: (2025)
Towards Personalized Evaluation of Large Language Models with An Anonymous Crowd-Sourcing Platform
di: Cheng, Mingyue, et al.
Pubblicazione: (2024)
di: Cheng, Mingyue, et al.
Pubblicazione: (2024)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
di: Xiao, Tong, et al.
Pubblicazione: (2025)
di: Xiao, Tong, et al.
Pubblicazione: (2025)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
di: Li, Chenglin, et al.
Pubblicazione: (2024)
di: Li, Chenglin, et al.
Pubblicazione: (2024)
Coupling Water Purification and Carbon Sequestration at Various Spatial Scales From Supply and Demand Perspective
di: Jing Cheng, et al.
Pubblicazione: (2026)
di: Jing Cheng, et al.
Pubblicazione: (2026)
WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
di: Li, Xin, et al.
Pubblicazione: (2025)
di: Li, Xin, et al.
Pubblicazione: (2025)
Learning Recommender Systems with Soft Target: A Decoupled Perspective
di: Zhang, Hao, et al.
Pubblicazione: (2024)
di: Zhang, Hao, et al.
Pubblicazione: (2024)
A Survey on Deep Text Hashing: Efficient Semantic Text Retrieval with Binary Representation
di: He, Liyang, et al.
Pubblicazione: (2025)
di: He, Liyang, et al.
Pubblicazione: (2025)
Empowering Sequential Recommendation from Collaborative Signals and Semantic Relatedness
di: Cheng, Mingyue, et al.
Pubblicazione: (2024)
di: Cheng, Mingyue, et al.
Pubblicazione: (2024)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
di: Zeng, Liang, et al.
Pubblicazione: (2024)
di: Zeng, Liang, et al.
Pubblicazione: (2024)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
EternalMath: A Living Benchmark of Frontier Mathematics that Evolves with Human Discovery
di: Ma, Jicheng, et al.
Pubblicazione: (2026)
di: Ma, Jicheng, et al.
Pubblicazione: (2026)
Hierarchical Multimodal LLMs with Semantic Space Alignment for Enhanced Time Series Classification
di: Tao, Xiaoyu, et al.
Pubblicazione: (2024)
di: Tao, Xiaoyu, et al.
Pubblicazione: (2024)
A Survey of Knowledge Tracing: Models, Variants, and Applications
di: Shen, Shuanghong, et al.
Pubblicazione: (2021)
di: Shen, Shuanghong, et al.
Pubblicazione: (2021)
Can Slow-thinking LLMs Reason Over Time? Empirical Studies in Time Series Forecasting
di: Cheng, Mingyue, et al.
Pubblicazione: (2025)
di: Cheng, Mingyue, et al.
Pubblicazione: (2025)
Documenti analoghi
-
UniCog: Uncovering Cognitive Abilities of LLMs through Latent Mind Space Analysis
di: Liu, Jiayu, et al.
Pubblicazione: (2026) -
Learning to Solve Geometry Problems via Simulating Human Dual-Reasoning Process
di: Xiao, Tong, et al.
Pubblicazione: (2024) -
Unified Uncertainty Estimation for Cognitive Diagnosis Models
di: Wang, Fei, et al.
Pubblicazione: (2024) -
Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
di: Huang, Zhenya, et al.
Pubblicazione: (2025) -
Bit-mask Robust Contrastive Knowledge Distillation for Unsupervised Semantic Hashing
di: He, Liyang, et al.
Pubblicazione: (2024)