Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhong, Qihuang, Wang, Kang, Xu, Ziyang, Liu, Juhua, Ding, Liang, Du, Bo |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation
par: Zhong, Qihuang, et autres
Publié: (2022)
par: Zhong, Qihuang, et autres
Publié: (2022)
Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
par: Zhong, Qihuang, et autres
Publié: (2026)
par: Zhong, Qihuang, et autres
Publié: (2026)
Can LLMs Solve longer Math Word Problems Better?
par: Xu, Xin, et autres
Publié: (2024)
par: Xu, Xin, et autres
Publié: (2024)
What Makes Math Word Problems Challenging for LLMs?
par: Srivatsa, KV Aditya, et autres
Publié: (2024)
par: Srivatsa, KV Aditya, et autres
Publié: (2024)
KaFT: Knowledge-aware Fine-tuning for Boosting LLMs' Domain-specific Question-Answering Performance
par: Zhong, Qihuang, et autres
Publié: (2025)
par: Zhong, Qihuang, et autres
Publié: (2025)
ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding
par: Zhong, Qihuang, et autres
Publié: (2024)
par: Zhong, Qihuang, et autres
Publié: (2024)
PANDA: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation
par: Zhong, Qihuang, et autres
Publié: (2022)
par: Zhong, Qihuang, et autres
Publié: (2022)
Try, Check and Retry: A Divide-and-Conquer Framework for Boosting Long-context Tool-Calling Performance of LLMs
par: Chen, Kunfeng, et autres
Publié: (2026)
par: Chen, Kunfeng, et autres
Publié: (2026)
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
par: Yuan, Fan, et autres
Publié: (2025)
par: Yuan, Fan, et autres
Publié: (2025)
MathDuels: Evaluating LLMs as Problem Posers and Solvers
par: Xu, Zhiqiu, et autres
Publié: (2026)
par: Xu, Zhiqiu, et autres
Publié: (2026)
Resolving Knowledge Conflicts in Domain-specific Data Selection: A Case Study on Medical Instruction-tuning
par: Zhong, Qihuang, et autres
Publié: (2025)
par: Zhong, Qihuang, et autres
Publié: (2025)
Revisiting Knowledge Distillation for Autoregressive Language Models
par: Zhong, Qihuang, et autres
Publié: (2024)
par: Zhong, Qihuang, et autres
Publié: (2024)
Learning from Imperfect Data: Towards Efficient Knowledge Distillation of Autoregressive Language Models for Text-to-SQL
par: Zhong, Qihuang, et autres
Publié: (2024)
par: Zhong, Qihuang, et autres
Publié: (2024)
GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
par: Li, Qintong, et autres
Publié: (2024)
par: Li, Qintong, et autres
Publié: (2024)
ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering
par: Zhu, Yikai, et autres
Publié: (2026)
par: Zhu, Yikai, et autres
Publié: (2026)
Iterative Data Generation with Large Language Models for Aspect-based Sentiment Analysis
par: Zhong, Qihuang, et autres
Publié: (2024)
par: Zhong, Qihuang, et autres
Publié: (2024)
Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
par: Qi, Zhenting, et autres
Publié: (2024)
par: Qi, Zhenting, et autres
Publié: (2024)
Adversarial Math Word Problem Generation
par: Xie, Roy, et autres
Publié: (2024)
par: Xie, Roy, et autres
Publié: (2024)
Expression Syntax Information Bottleneck for Math Word Problems
par: Xiong, Jing, et autres
Publié: (2023)
par: Xiong, Jing, et autres
Publié: (2023)
Self-consistent Reasoning For Solving Math Word Problems
par: Xiong, Jing, et autres
Publié: (2022)
par: Xiong, Jing, et autres
Publié: (2022)
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems
par: Kai, Ding, et autres
Publié: (2024)
par: Kai, Ding, et autres
Publié: (2024)
Template-Driven LLM-Paraphrased Framework for Tabular Math Word Problem Generation
par: Kang, Xiaoqiang, et autres
Publié: (2024)
par: Kang, Xiaoqiang, et autres
Publié: (2024)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
par: Ye, Maoyuan, et autres
Publié: (2025)
par: Ye, Maoyuan, et autres
Publié: (2025)
We Need Knowledge Distillation for Solving Math Word Problems
par: Shen, Zhenquan, et autres
Publié: (2025)
par: Shen, Zhenquan, et autres
Publié: (2025)
Structured Reasoning with Tree-of-Thoughts for Bengali Math Word Problems
par: Mahmood, Aurprita, et autres
Publié: (2025)
par: Mahmood, Aurprita, et autres
Publié: (2025)
EDUMATH: Generating Standards-aligned Educational Math Word Problems
par: Christ, Bryan R., et autres
Publié: (2025)
par: Christ, Bryan R., et autres
Publié: (2025)
MATHWELL: Generating Educational Math Word Problems Using Teacher Annotations
par: Christ, Bryan R, et autres
Publié: (2024)
par: Christ, Bryan R, et autres
Publié: (2024)
Elementary Math Word Problem Generation using Large Language Models
par: Ariyarathne, Nimesh, et autres
Publié: (2025)
par: Ariyarathne, Nimesh, et autres
Publié: (2025)
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
par: Anantheswaran, Ujjwala, et autres
Publié: (2024)
par: Anantheswaran, Ujjwala, et autres
Publié: (2024)
Augmenting Math Word Problems via Iterative Question Composing
par: Liu, Haoxiong, et autres
Publié: (2024)
par: Liu, Haoxiong, et autres
Publié: (2024)
From Large to Tiny: Distilling and Refining Mathematical Expertise for Math Word Problems with Weakly Supervision
par: Lin, Qingwen, et autres
Publié: (2024)
par: Lin, Qingwen, et autres
Publié: (2024)
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
par: Du, Yexing, et autres
Publié: (2024)
par: Du, Yexing, et autres
Publié: (2024)
Solving Math Word Problems via Cooperative Reasoning induced Language Models
par: Zhu, Xinyu, et autres
Publié: (2022)
par: Zhu, Xinyu, et autres
Publié: (2022)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
par: Sun, Yuhong, et autres
Publié: (2024)
par: Sun, Yuhong, et autres
Publié: (2024)
Data Augmentation with In-Context Learning and Comparative Evaluation in Math Word Problem Solving
par: Yigit, Gulsum, et autres
Publié: (2024)
par: Yigit, Gulsum, et autres
Publié: (2024)
Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems
par: Yang, Kaiqi, et autres
Publié: (2025)
par: Yang, Kaiqi, et autres
Publié: (2025)
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
par: Zhang, Yingji, et autres
Publié: (2026)
par: Zhang, Yingji, et autres
Publié: (2026)
LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems
par: Xu, Nan, et autres
Publié: (2024)
par: Xu, Nan, et autres
Publié: (2024)
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
par: Cheng, Ziling, et autres
Publié: (2025)
par: Cheng, Ziling, et autres
Publié: (2025)
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework
par: Sun, Yuhong, et autres
Publié: (2025)
par: Sun, Yuhong, et autres
Publié: (2025)
Documents similaires
-
E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation
par: Zhong, Qihuang, et autres
Publié: (2022) -
Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
par: Zhong, Qihuang, et autres
Publié: (2026) -
Can LLMs Solve longer Math Word Problems Better?
par: Xu, Xin, et autres
Publié: (2024) -
What Makes Math Word Problems Challenging for LLMs?
par: Srivatsa, KV Aditya, et autres
Publié: (2024) -
KaFT: Knowledge-aware Fine-tuning for Boosting LLMs' Domain-specific Question-Answering Performance
par: Zhong, Qihuang, et autres
Publié: (2025)