Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Yanqi, Ji, Yuxiang, Zhang, Xiao, Wang, Yong, Chu, Xiangxiang, Lu, Zhiwu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty
by: Dai, Yanqi, et al.
Published: (2024)
by: Dai, Yanqi, et al.
Published: (2024)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
by: Dipta, Shubhashis Roy, et al.
Published: (2026)
by: Dipta, Shubhashis Roy, et al.
Published: (2026)
HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation
by: Xiong, Feng, et al.
Published: (2025)
by: Xiong, Feng, et al.
Published: (2025)
D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning
by: Zhang, Ru, et al.
Published: (2026)
by: Zhang, Ru, et al.
Published: (2026)
GRPO-LEAD: A Difficulty-Aware Reinforcement Learning Approach for Concise Mathematical Reasoning in Language Models
by: Zhang, Jixiao, et al.
Published: (2025)
by: Zhang, Jixiao, et al.
Published: (2025)
NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty
by: Zotos, Leonidas, et al.
Published: (2025)
by: Zotos, Leonidas, et al.
Published: (2025)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
by: Tong, Yuxuan, et al.
Published: (2024)
by: Tong, Yuxuan, et al.
Published: (2024)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
by: Hong, Zijin, et al.
Published: (2025)
by: Hong, Zijin, et al.
Published: (2025)
AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting
by: Huang, Shijue, et al.
Published: (2025)
by: Huang, Shijue, et al.
Published: (2025)
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
by: Ma, Ziyu, et al.
Published: (2026)
by: Ma, Ziyu, et al.
Published: (2026)
Conversational Question Answering with Reformulations over Knowledge Graph
by: Liu, Lihui, et al.
Published: (2023)
by: Liu, Lihui, et al.
Published: (2023)
I Could've Asked That: Reformulating Unanswerable Questions
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
by: Zhang, Xiaoying, et al.
Published: (2025)
by: Zhang, Xiaoying, et al.
Published: (2025)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
by: Lu, Zimu, et al.
Published: (2024)
by: Lu, Zimu, et al.
Published: (2024)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
by: Xiao, Wenyi, et al.
Published: (2025)
by: Xiao, Wenyi, et al.
Published: (2025)
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
by: Zhu, Yubo, et al.
Published: (2025)
by: Zhu, Yubo, et al.
Published: (2025)
Preference Alignment for Diffusion Model via Explicit Denoised Distribution Estimation
by: Shi, Dingyuan, et al.
Published: (2024)
by: Shi, Dingyuan, et al.
Published: (2024)
ExGRPO: Learning to Reason from Experience
by: Zhan, Runzhe, et al.
Published: (2025)
by: Zhan, Runzhe, et al.
Published: (2025)
Robust Training for Conversational Question Answering Models with Reinforced Reformulation Generation
by: Kaiser, Magdalena, et al.
Published: (2023)
by: Kaiser, Magdalena, et al.
Published: (2023)
BoostTaxo: Zero-Shot Taxonomy Induction via Boosting-Style Agentic Reasoning and Constraint-Aware Calibration
by: Ling, Yancheng, et al.
Published: (2026)
by: Ling, Yancheng, et al.
Published: (2026)
Temporal-Aware Heterogeneous Graph Reasoning with Multi-View Fusion for Temporal Question Answering
by: Wen, Wuzhenghong, et al.
Published: (2026)
by: Wen, Wuzhenghong, et al.
Published: (2026)
RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty
by: Zhang, Ziqian, et al.
Published: (2026)
by: Zhang, Ziqian, et al.
Published: (2026)
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
by: Shao, Zhihong, et al.
Published: (2025)
by: Shao, Zhihong, et al.
Published: (2025)
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
by: Zheng, Jinquan, et al.
Published: (2026)
by: Zheng, Jinquan, et al.
Published: (2026)
StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation
by: Jang, Geonhui, et al.
Published: (2026)
by: Jang, Geonhui, et al.
Published: (2026)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
by: Wei, Kangda, et al.
Published: (2026)
by: Wei, Kangda, et al.
Published: (2026)
The Harder The Better: Maintaining Supervised Fine-tuning Generalization with Less but Harder Data
by: Shang, Zhaoyang, et al.
Published: (2025)
by: Shang, Zhaoyang, et al.
Published: (2025)
Towards Better Generalization in Open-Domain Question Answering by Mitigating Context Memorization
by: Zhang, Zixuan, et al.
Published: (2024)
by: Zhang, Zixuan, et al.
Published: (2024)
From Reasoning to Code: GRPO Optimization for Underrepresented Languages
by: Pennino, Federico, et al.
Published: (2025)
by: Pennino, Federico, et al.
Published: (2025)
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
by: Ji, Yuxiang, et al.
Published: (2026)
by: Ji, Yuxiang, et al.
Published: (2026)
Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions
by: Jing, Dong, et al.
Published: (2025)
by: Jing, Dong, et al.
Published: (2025)
The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
by: Tan, Yuqiao, et al.
Published: (2025)
by: Tan, Yuqiao, et al.
Published: (2025)
Omne-R1: Learning to Reason with Memory for Multi-hop Question Answering
by: Liu, Boyuan, et al.
Published: (2025)
by: Liu, Boyuan, et al.
Published: (2025)
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
by: Do, Heejin, et al.
Published: (2025)
by: Do, Heejin, et al.
Published: (2025)
BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question Answering
by: Chu, Zheng, et al.
Published: (2024)
by: Chu, Zheng, et al.
Published: (2024)
Outdated Issue Aware Decoding for Reasoning Questions on Edited Knowledge
by: Sun, Zengkui, et al.
Published: (2024)
by: Sun, Zengkui, et al.
Published: (2024)
Similar Items
-
Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty
by: Dai, Yanqi, et al.
Published: (2024) -
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
by: Dipta, Shubhashis Roy, et al.
Published: (2026) -
HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation
by: Xiong, Feng, et al.
Published: (2025) -
D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learning
by: Zhang, Ru, et al.
Published: (2026) -
GRPO-LEAD: A Difficulty-Aware Reinforcement Learning Approach for Concise Mathematical Reasoning in Language Models
by: Zhang, Jixiao, et al.
Published: (2025)