RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xinyuan, Xu, Murong, Tao, Wenbiao, Zhu, Hanlun, Zhao, Yike, Zhang, Jipeng, Lan, Yunshi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TagRAG: Tag-guided Hierarchical Knowledge Graph Retrieval-Augmented Generation
by: Tao, Wenbiao, et al.
Published: (2025)
by: Tao, Wenbiao, et al.
Published: (2025)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025)
by: Schmucker, Robin, et al.
Published: (2025)
Unsupervised Text Style Transfer for Controllable Intensity
by: Gu, Shuhuan, et al.
Published: (2026)
by: Gu, Shuhuan, et al.
Published: (2026)
ComRAG: Retrieval-Augmented Generation with Dynamic Vector Stores for Real-time Community Question Answering in Industry
by: Chen, Qinwen, et al.
Published: (2025)
by: Chen, Qinwen, et al.
Published: (2025)
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
by: Cong, Longwei, et al.
Published: (2026)
by: Cong, Longwei, et al.
Published: (2026)
Aligning Large Language Models to a Domain-specific Graph Database for NL2GQL
by: Liang, Yuanyuan, et al.
Published: (2024)
by: Liang, Yuanyuan, et al.
Published: (2024)
Prediction of Item Difficulty for Reading Comprehension Items by Creation of Annotated Item Repository
by: Kapoor, Radhika, et al.
Published: (2025)
by: Kapoor, Radhika, et al.
Published: (2025)
TreeEval: Benchmark-Free Evaluation of Large Language Models through Tree Planning
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
EvolveSearch: An Iterative Self-Evolving Search Agent
by: Zhang, Dingchu, et al.
Published: (2025)
by: Zhang, Dingchu, et al.
Published: (2025)
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
by: Zhou, Hongli, et al.
Published: (2025)
by: Zhou, Hongli, et al.
Published: (2025)
Scoring Edit Impact in Grammatical Error Correction via Embedded Association Graphs
by: Xiao, Qiyuan, et al.
Published: (2026)
by: Xiao, Qiyuan, et al.
Published: (2026)
Auditing LLM Benchmarks with Item Response Theory
by: Land, Sander, et al.
Published: (2026)
by: Land, Sander, et al.
Published: (2026)
Survey of Natural Language Processing for Education: Taxonomy, Systematic Review, and Future Trends
by: Lan, Yunshi, et al.
Published: (2024)
by: Lan, Yunshi, et al.
Published: (2024)
MathVC: An LLM-Simulated Multi-Character Virtual Classroom for Mathematics Education
by: Yue, Murong, et al.
Published: (2024)
by: Yue, Murong, et al.
Published: (2024)
GRPO-LEAD: A Difficulty-Aware Reinforcement Learning Approach for Concise Mathematical Reasoning in Language Models
by: Zhang, Jixiao, et al.
Published: (2025)
by: Zhang, Jixiao, et al.
Published: (2025)
A Survey of Large Language Model Agents for Question Answering
by: Yue, Murong
Published: (2025)
by: Yue, Murong
Published: (2025)
More Data or Better Data? A Critical Analysis of Data Selection and Synthesis for Mathematical Reasoning
by: Zhao, Yike, et al.
Published: (2025)
by: Zhao, Yike, et al.
Published: (2025)
Unsupervised Text Style Transfer via LLMs and Attention Masking with Multi-way Interactions
by: Pan, Lei, et al.
Published: (2024)
by: Pan, Lei, et al.
Published: (2024)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory
by: Uebayashi, Shunki, et al.
Published: (2026)
by: Uebayashi, Shunki, et al.
Published: (2026)
Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning
by: Yue, Murong, et al.
Published: (2023)
by: Yue, Murong, et al.
Published: (2023)
Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation
by: Dai, Yanqi, et al.
Published: (2026)
by: Dai, Yanqi, et al.
Published: (2026)
Unleashing the Power of Large Language Models in Zero-shot Relation Extraction via Self-Prompting
by: Liu, Siyi, et al.
Published: (2024)
by: Liu, Siyi, et al.
Published: (2024)
Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning
by: Zhang, Lan, et al.
Published: (2025)
by: Zhang, Lan, et al.
Published: (2025)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
Can Model Uncertainty Function as a Proxy for Multiple-Choice Question Item Difficulty?
by: Zotos, Leonidas, et al.
Published: (2024)
by: Zotos, Leonidas, et al.
Published: (2024)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
by: Dipta, Shubhashis Roy, et al.
Published: (2026)
by: Dipta, Shubhashis Roy, et al.
Published: (2026)
An LLM-Enhanced Adversarial Editing System for Lexical Simplification
by: Tan, Keren, et al.
Published: (2024)
by: Tan, Keren, et al.
Published: (2024)
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment
by: Zhang, Jingshen, et al.
Published: (2024)
by: Zhang, Jingshen, et al.
Published: (2024)
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
by: Cao, Chuxue, et al.
Published: (2025)
by: Cao, Chuxue, et al.
Published: (2025)
ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
by: Liu, Hongwei, et al.
Published: (2025)
by: Liu, Hongwei, et al.
Published: (2025)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
by: Acquaye, Christabel, et al.
Published: (2026)
by: Acquaye, Christabel, et al.
Published: (2026)
A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation
by: Hwang, Seonjeong, et al.
Published: (2026)
by: Hwang, Seonjeong, et al.
Published: (2026)
PACE: Prefix-Protected and Difficulty-Aware Compression for Efficient Reasoning
by: Feng, Ruixiang, et al.
Published: (2026)
by: Feng, Ruixiang, et al.
Published: (2026)
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
by: Zhang, Jiaqiao, et al.
Published: (2026)
by: Zhang, Jiaqiao, et al.
Published: (2026)
CODA: Difficulty-Aware Compute Allocation for Adaptive Reasoning
by: Wu, Siye, et al.
Published: (2026)
by: Wu, Siye, et al.
Published: (2026)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
by: Tong, Yuxuan, et al.
Published: (2024)
by: Tong, Yuxuan, et al.
Published: (2024)
Chain-of-Tools: Utilizing Massive Unseen Tools in the CoT Reasoning of Frozen Language Models
by: Wu, Mengsong, et al.
Published: (2025)
by: Wu, Mengsong, et al.
Published: (2025)
Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News Detection
by: Zhang, Chaowei, et al.
Published: (2025)
by: Zhang, Chaowei, et al.
Published: (2025)
Similar Items
-
TagRAG: Tag-guided Hierarchical Knowledge Graph Retrieval-Augmented Generation
by: Tao, Wenbiao, et al.
Published: (2025) -
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025) -
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025) -
Unsupervised Text Style Transfer for Controllable Intensity
by: Gu, Shuhuan, et al.
Published: (2026) -
ComRAG: Retrieval-Augmented Generation with Dynamic Vector Stores for Real-time Community Question Answering in Industry
by: Chen, Qinwen, et al.
Published: (2025)