Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis
Fuente:
arXiv
Guardado en:
| Autores principales: | Fu, Jiayu, Heddaya, Mourad, Tan, Chenhao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Language of Bargaining
por: Heddaya, Mourad, et al.
Publicado: (2023)
por: Heddaya, Mourad, et al.
Publicado: (2023)
Causal Micro-Narratives
por: Heddaya, Mourad, et al.
Publicado: (2024)
por: Heddaya, Mourad, et al.
Publicado: (2024)
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions
por: Heddaya, Mourad, et al.
Publicado: (2024)
por: Heddaya, Mourad, et al.
Publicado: (2024)
LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators
por: Shah, Cheril, et al.
Publicado: (2025)
por: Shah, Cheril, et al.
Publicado: (2025)
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
por: Li, Mingxuan, et al.
Publicado: (2025)
por: Li, Mingxuan, et al.
Publicado: (2025)
Hypothesis Generation with Large Language Models
por: Zhou, Yangqiaoyu, et al.
Publicado: (2024)
por: Zhou, Yangqiaoyu, et al.
Publicado: (2024)
Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
por: Huang, Zhenya, et al.
Publicado: (2025)
por: Huang, Zhenya, et al.
Publicado: (2025)
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
por: Liu, Haokun, et al.
Publicado: (2025)
por: Liu, Haokun, et al.
Publicado: (2025)
Literature Meets Data: A Synergistic Approach to Hypothesis Generation
por: Liu, Haokun, et al.
Publicado: (2024)
por: Liu, Haokun, et al.
Publicado: (2024)
Adversarial Math Word Problem Generation
por: Xie, Roy, et al.
Publicado: (2024)
por: Xie, Roy, et al.
Publicado: (2024)
Solving Math Word Problems Using Estimation Verification and Equation Generation
por: Piehl, Mitchell, et al.
Publicado: (2025)
por: Piehl, Mitchell, et al.
Publicado: (2025)
HardCore Generation: Generating Hard UNSAT Problems for Data Augmentation
por: Cotnareanu, Joseph, et al.
Publicado: (2024)
por: Cotnareanu, Joseph, et al.
Publicado: (2024)
DiagramIR: An Automatic Pipeline for Educational Math Diagram Evaluation
por: Kumar, Vishal, et al.
Publicado: (2025)
por: Kumar, Vishal, et al.
Publicado: (2025)
EDUMATH: Generating Standards-aligned Educational Math Word Problems
por: Christ, Bryan R., et al.
Publicado: (2025)
por: Christ, Bryan R., et al.
Publicado: (2025)
LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
por: Pan, Zhuoshi, et al.
Publicado: (2025)
por: Pan, Zhuoshi, et al.
Publicado: (2025)
Agentic Application in Power Grid Static Analysis: Automatic Code Generation and Error Correction
por: Wang, Qinjuan, et al.
Publicado: (2026)
por: Wang, Qinjuan, et al.
Publicado: (2026)
Towards more Contextual Agents: An extractor-Generator Optimization Framework
por: Aouini, Mourad, et al.
Publicado: (2025)
por: Aouini, Mourad, et al.
Publicado: (2025)
Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
por: Guo, Zixian, et al.
Publicado: (2025)
por: Guo, Zixian, et al.
Publicado: (2025)
MathMistake Checker: A Comprehensive Demonstration for Step-by-Step Math Problem Mistake Finding by Prompt-Guided LLMs
por: Zhang, Tianyang, et al.
Publicado: (2025)
por: Zhang, Tianyang, et al.
Publicado: (2025)
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
por: Nguyen, Dang, et al.
Publicado: (2025)
por: Nguyen, Dang, et al.
Publicado: (2025)
Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models
por: Kim, Hyunwoo, et al.
Publicado: (2025)
por: Kim, Hyunwoo, et al.
Publicado: (2025)
Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems
por: Pan, Xi-Wei, et al.
Publicado: (2026)
por: Pan, Xi-Wei, et al.
Publicado: (2026)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
por: Khan, Zaid, et al.
Publicado: (2025)
por: Khan, Zaid, et al.
Publicado: (2025)
Controllable Logical Hypothesis Generation for Abductive Reasoning in Knowledge Graphs
por: Gao, Yisen, et al.
Publicado: (2025)
por: Gao, Yisen, et al.
Publicado: (2025)
On Size and Hardness Generalization in Unsupervised Learning for the Travelling Salesman Problem
por: Min, Yimeng, et al.
Publicado: (2024)
por: Min, Yimeng, et al.
Publicado: (2024)
Decomposing LLM Self-Correction: The Accuracy-Correction Paradox and Error Depth Hypothesis
por: Li, Yin
Publicado: (2025)
por: Li, Yin
Publicado: (2025)
LLM Agent Swarm for Hypothesis-Driven Drug Discovery
por: Song, Kevin, et al.
Publicado: (2025)
por: Song, Kevin, et al.
Publicado: (2025)
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
por: Liu, Jiayu, et al.
Publicado: (2025)
por: Liu, Jiayu, et al.
Publicado: (2025)
Neuro-Symbolic Data Generation for Math Reasoning
por: Li, Zenan, et al.
Publicado: (2024)
por: Li, Zenan, et al.
Publicado: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
por: Fang, Meng, et al.
Publicado: (2024)
por: Fang, Meng, et al.
Publicado: (2024)
IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
por: Biyani, Param, et al.
Publicado: (2025)
por: Biyani, Param, et al.
Publicado: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
por: Huang, Kaixuan, et al.
Publicado: (2025)
por: Huang, Kaixuan, et al.
Publicado: (2025)
StepMathAgent: A Step-Wise Agent for Evaluating Mathematical Processes through Tree-of-Error
por: Yang, Shu-Xun, et al.
Publicado: (2025)
por: Yang, Shu-Xun, et al.
Publicado: (2025)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
por: Lai, Yuhang, et al.
Publicado: (2026)
por: Lai, Yuhang, et al.
Publicado: (2026)
Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models
por: Wang, Junling, et al.
Publicado: (2025)
por: Wang, Junling, et al.
Publicado: (2025)
The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?
por: Sourav, Shashwat, et al.
Publicado: (2026)
por: Sourav, Shashwat, et al.
Publicado: (2026)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
por: Liu, Xianyang, et al.
Publicado: (2025)
por: Liu, Xianyang, et al.
Publicado: (2025)
Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling
por: Tiblias, Federico, et al.
Publicado: (2025)
por: Tiblias, Federico, et al.
Publicado: (2025)
Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
por: Mahran, Mariam, et al.
Publicado: (2025)
por: Mahran, Mariam, et al.
Publicado: (2025)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
por: Tong, Yuxuan, et al.
Publicado: (2024)
por: Tong, Yuxuan, et al.
Publicado: (2024)
Ejemplares similares
-
Language of Bargaining
por: Heddaya, Mourad, et al.
Publicado: (2023) -
Causal Micro-Narratives
por: Heddaya, Mourad, et al.
Publicado: (2024) -
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions
por: Heddaya, Mourad, et al.
Publicado: (2024) -
LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators
por: Shah, Cheril, et al.
Publicado: (2025) -
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
por: Li, Mingxuan, et al.
Publicado: (2025)