Large Language Models Struggle with Unreasonability in Math Problems
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Jingyuan, Dai, Damai, Yuan, Zihang, li, Rui, Luo, Weilin, Wang, Bin, Liu, Qun, Sha, Lei, Sui, Zhifang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Language Models Encode the Value of Numbers Linearly
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
Exploring Activation Patterns of Parameters in Language Models
di: Wang, Yudong, et al.
Pubblicazione: (2024)
di: Wang, Yudong, et al.
Pubblicazione: (2024)
Plug-and-Play Training Framework for Preference Optimization
di: Ma, Jingyuan, et al.
Pubblicazione: (2024)
di: Ma, Jingyuan, et al.
Pubblicazione: (2024)
Not All Demonstration Examples are Equally Beneficial: Reweighting Demonstration Examples for In-Context Learning
di: Yang, Zhe, et al.
Pubblicazione: (2023)
di: Yang, Zhe, et al.
Pubblicazione: (2023)
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
di: Wang, Peiyi, et al.
Pubblicazione: (2023)
di: Wang, Peiyi, et al.
Pubblicazione: (2023)
Towards Harmonized Uncertainty Estimation for Large Language Models
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
HauntAttack: When Attack Follows Reasoning as a Shadow
di: Ma, Jingyuan, et al.
Pubblicazione: (2025)
di: Ma, Jingyuan, et al.
Pubblicazione: (2025)
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
di: An, Shengnan, et al.
Pubblicazione: (2025)
di: An, Shengnan, et al.
Pubblicazione: (2025)
A Survey on In-context Learning
di: Dong, Qingxiu, et al.
Pubblicazione: (2022)
di: Dong, Qingxiu, et al.
Pubblicazione: (2022)
Taking a Deep Breath: Enhancing Language Modeling of Large Language Models with Sentinel Tokens
di: Luo, Weiyao, et al.
Pubblicazione: (2024)
di: Luo, Weiyao, et al.
Pubblicazione: (2024)
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?
di: Yang, Zhe, et al.
Pubblicazione: (2024)
di: Yang, Zhe, et al.
Pubblicazione: (2024)
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
di: Zhan, Weidong, et al.
Pubblicazione: (2025)
di: Zhan, Weidong, et al.
Pubblicazione: (2025)
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
di: Xue, Boyang, et al.
Pubblicazione: (2025)
di: Xue, Boyang, et al.
Pubblicazione: (2025)
PeriodicLoRA: Breaking the Low-Rank Bottleneck in LoRA Optimization
di: Meng, Xiangdi, et al.
Pubblicazione: (2024)
di: Meng, Xiangdi, et al.
Pubblicazione: (2024)
LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems
di: Xu, Nan, et al.
Pubblicazione: (2024)
di: Xu, Nan, et al.
Pubblicazione: (2024)
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems
di: Kai, Ding, et al.
Pubblicazione: (2024)
di: Kai, Ding, et al.
Pubblicazione: (2024)
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
di: Xu, Yifan, et al.
Pubblicazione: (2024)
di: Xu, Yifan, et al.
Pubblicazione: (2024)
Self-Boosting Large Language Models with Synthetic Preference Data
di: Dong, Qingxiu, et al.
Pubblicazione: (2024)
di: Dong, Qingxiu, et al.
Pubblicazione: (2024)
Elementary Math Word Problem Generation using Large Language Models
di: Ariyarathne, Nimesh, et al.
Pubblicazione: (2025)
di: Ariyarathne, Nimesh, et al.
Pubblicazione: (2025)
MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model
di: Yang, Zhen, et al.
Pubblicazione: (2024)
di: Yang, Zhen, et al.
Pubblicazione: (2024)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
di: Li, Zheng, et al.
Pubblicazione: (2025)
di: Li, Zheng, et al.
Pubblicazione: (2025)
MathLearner: A Large Language Model Agent Framework for Learning to Solve Mathematical Problems
di: Xie, Wenbei, et al.
Pubblicazione: (2024)
di: Xie, Wenbei, et al.
Pubblicazione: (2024)
CoLT: Reasoning with Chain of Latent Tool Calls
di: Zhu, Fangwei, et al.
Pubblicazione: (2026)
di: Zhu, Fangwei, et al.
Pubblicazione: (2026)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
di: Sun, Yuhong, et al.
Pubblicazione: (2024)
di: Sun, Yuhong, et al.
Pubblicazione: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
di: Ying, Huaiyuan, et al.
Pubblicazione: (2024)
di: Ying, Huaiyuan, et al.
Pubblicazione: (2024)
Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models
di: Wang, Zihan, et al.
Pubblicazione: (2024)
di: Wang, Zihan, et al.
Pubblicazione: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
di: Fang, Meng, et al.
Pubblicazione: (2024)
di: Fang, Meng, et al.
Pubblicazione: (2024)
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
di: Zhang, Yingji, et al.
Pubblicazione: (2026)
di: Zhang, Yingji, et al.
Pubblicazione: (2026)
Improving Math Problem Solving in Large Language Models Through Categorization and Strategy Tailoring
di: Akella, Amogh
Pubblicazione: (2024)
di: Akella, Amogh
Pubblicazione: (2024)
Large Language Models Struggle in Token-Level Clinical Named Entity Recognition
di: Lu, Qiuhao, et al.
Pubblicazione: (2024)
di: Lu, Qiuhao, et al.
Pubblicazione: (2024)
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
di: Yang, Yixin, et al.
Pubblicazione: (2024)
di: Yang, Yixin, et al.
Pubblicazione: (2024)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
di: Shi, Wenhao, et al.
Pubblicazione: (2024)
di: Shi, Wenhao, et al.
Pubblicazione: (2024)
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
di: Dai, Damai, et al.
Pubblicazione: (2024)
di: Dai, Damai, et al.
Pubblicazione: (2024)
The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
di: Lucy, Li, et al.
Pubblicazione: (2026)
di: Lucy, Li, et al.
Pubblicazione: (2026)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
di: Sun, Yuhong, et al.
Pubblicazione: (2025)
Prompt and Parameter Co-Optimization for Large Language Models
di: Bo, Xiaohe, et al.
Pubblicazione: (2025)
di: Bo, Xiaohe, et al.
Pubblicazione: (2025)
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
di: Colle, Vincenzo, et al.
Pubblicazione: (2025)
di: Colle, Vincenzo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Language Models Encode the Value of Numbers Linearly
di: Zhu, Fangwei, et al.
Pubblicazione: (2024) -
Exploring Activation Patterns of Parameters in Language Models
di: Wang, Yudong, et al.
Pubblicazione: (2024) -
Plug-and-Play Training Framework for Preference Optimization
di: Ma, Jingyuan, et al.
Pubblicazione: (2024) -
Not All Demonstration Examples are Equally Beneficial: Reweighting Demonstration Examples for In-Context Learning
di: Yang, Zhe, et al.
Pubblicazione: (2023) -
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
di: Li, Rui, et al.
Pubblicazione: (2025)