A Diversity-Enhanced Knowledge Distillation Model for Practical Math Word Problem Solving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909450473832448 |
|---|---|
| author | Zhang, Yi Zhou, Guangyou Xie, Zhiwen Ma, Jinjin Huang, Jimmy Xiangji |
| author_facet | Zhang, Yi Zhou, Guangyou Xie, Zhiwen Ma, Jinjin Huang, Jimmy Xiangji |
| contents | Math Word Problem (MWP) solving is a critical task in natural language processing, has garnered significant research interest in recent years. Various recent studies heavily rely on Seq2Seq models and their extensions (e.g., Seq2Tree and Graph2Tree) to generate mathematical equations. While effective, these models struggle to generate diverse but counterpart solution equations, limiting their generalization across various math problem scenarios. In this paper, we introduce a novel Diversity-enhanced Knowledge Distillation (DivKD) model for practical MWP solving. Our approach proposes an adaptive diversity distillation method, in which a student model learns diverse equations by selectively transferring high-quality knowledge from a teacher model. Additionally, we design a diversity prior-enhanced student model to better capture the diversity distribution of equations by incorporating a conditional variational auto-encoder. Extensive experiments on {four} MWP benchmark datasets demonstrate that our approach achieves higher answer accuracy than strong baselines while maintaining high efficiency for practical applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_03670 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Diversity-Enhanced Knowledge Distillation Model for Practical Math Word Problem Solving Zhang, Yi Zhou, Guangyou Xie, Zhiwen Ma, Jinjin Huang, Jimmy Xiangji Computation and Language Artificial Intelligence Math Word Problem (MWP) solving is a critical task in natural language processing, has garnered significant research interest in recent years. Various recent studies heavily rely on Seq2Seq models and their extensions (e.g., Seq2Tree and Graph2Tree) to generate mathematical equations. While effective, these models struggle to generate diverse but counterpart solution equations, limiting their generalization across various math problem scenarios. In this paper, we introduce a novel Diversity-enhanced Knowledge Distillation (DivKD) model for practical MWP solving. Our approach proposes an adaptive diversity distillation method, in which a student model learns diverse equations by selectively transferring high-quality knowledge from a teacher model. Additionally, we design a diversity prior-enhanced student model to better capture the diversity distribution of equations by incorporating a conditional variational auto-encoder. Extensive experiments on {four} MWP benchmark datasets demonstrate that our approach achieves higher answer accuracy than strong baselines while maintaining high efficiency for practical applications. |
| title | A Diversity-Enhanced Knowledge Distillation Model for Practical Math Word Problem Solving |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2501.03670 |