SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Manem, Chaitanya, Brahma, Pratik Prabhanjan, Mishra, Prakamya, Liu, Zicheng, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
von: Joshi, Vinay, et al.
Veröffentlicht: (2025)
von: Joshi, Vinay, et al.
Veröffentlicht: (2025)
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
von: Ray, Pretam, et al.
Veröffentlicht: (2026)
von: Ray, Pretam, et al.
Veröffentlicht: (2026)
Instella: Fully Open Language Models with Stellar Performance
von: Liu, Jiang, et al.
Veröffentlicht: (2025)
von: Liu, Jiang, et al.
Veröffentlicht: (2025)
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
von: Mishra, Prakamya, et al.
Veröffentlicht: (2025)
von: Mishra, Prakamya, et al.
Veröffentlicht: (2025)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
von: Wang, Jianghui, et al.
Veröffentlicht: (2025)
von: Wang, Jianghui, et al.
Veröffentlicht: (2025)
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
von: Singh, Shivam, et al.
Veröffentlicht: (2026)
ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2026)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2026)
Estimating the Usefulness of Clarifying Questions and Answers for Conversational Search
von: Sekulić, Ivan, et al.
Veröffentlicht: (2024)
von: Sekulić, Ivan, et al.
Veröffentlicht: (2024)
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
ExpertQA: Expert-Curated Questions and Attributed Answers
von: Malaviya, Chaitanya, et al.
Veröffentlicht: (2023)
von: Malaviya, Chaitanya, et al.
Veröffentlicht: (2023)
LLMs Provide Unstable Answers to Legal Questions
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
Automate Knowledge Concept Tagging on Math Questions with LLMs
von: Li, Hang, et al.
Veröffentlicht: (2024)
von: Li, Hang, et al.
Veröffentlicht: (2024)
Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs
von: Gupta, Himanshu, et al.
Veröffentlicht: (2026)
von: Gupta, Himanshu, et al.
Veröffentlicht: (2026)
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
von: Yu, Longhui, et al.
Veröffentlicht: (2023)
von: Yu, Longhui, et al.
Veröffentlicht: (2023)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
QueST: Incentivizing LLMs to Generate Difficult Problems
von: Hu, Hanxu, et al.
Veröffentlicht: (2025)
von: Hu, Hanxu, et al.
Veröffentlicht: (2025)
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
von: Zhu, Yunqi, et al.
Veröffentlicht: (2024)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Hybrid Graphs for Table-and-Text based Question Answering using LLMs
von: Agarwal, Ankush, et al.
Veröffentlicht: (2025)
von: Agarwal, Ankush, et al.
Veröffentlicht: (2025)
Putting People in LLMs' Shoes: Generating Better Answers via Question Rewriter
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
von: Chen, Junhao, et al.
Veröffentlicht: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
Agent Laboratory: Using LLM Agents as Research Assistants
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2025)
von: Schmidgall, Samuel, et al.
Veröffentlicht: (2025)
LLMs Encode How Difficult Problems Are
von: Lugoloobi, William, et al.
Veröffentlicht: (2025)
von: Lugoloobi, William, et al.
Veröffentlicht: (2025)
Do LLMs and Humans Find the Same Questions Difficult? A Case Study on Japanese Quiz Answering
von: Sugiura, Naoya, et al.
Veröffentlicht: (2025)
von: Sugiura, Naoya, et al.
Veröffentlicht: (2025)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
von: Van Long, Phuoc Pham, et al.
Veröffentlicht: (2023)
von: Van Long, Phuoc Pham, et al.
Veröffentlicht: (2023)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
CaptionQA: Is Your Caption as Useful as the Image Itself?
von: Yang, Shijia, et al.
Veröffentlicht: (2025)
von: Yang, Shijia, et al.
Veröffentlicht: (2025)
Question Answering with LLMs and Learning from Answer Sets
von: Borroto, Manuel, et al.
Veröffentlicht: (2025)
von: Borroto, Manuel, et al.
Veröffentlicht: (2025)
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset
von: Henkel, Owen, et al.
Veröffentlicht: (2023)
von: Henkel, Owen, et al.
Veröffentlicht: (2023)
Automatic Question & Answer Generation Using Generative Large Language Model (LLM)
von: Ehsan, Md. Alvee, et al.
Veröffentlicht: (2025)
von: Ehsan, Md. Alvee, et al.
Veröffentlicht: (2025)
Synergizing LLMs and Knowledge Graphs: A Novel Approach to Software Repository-Related Question Answering
von: Abedu, Samuel, et al.
Veröffentlicht: (2024)
von: Abedu, Samuel, et al.
Veröffentlicht: (2024)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
Automatic Question-Answer Generation for Long-Tail Knowledge
von: Kumar, Rohan, et al.
Veröffentlicht: (2024)
von: Kumar, Rohan, et al.
Veröffentlicht: (2024)
CoinMath: Harnessing the Power of Coding Instruction for Math LLMs
von: Wei, Chengwei, et al.
Veröffentlicht: (2024)
von: Wei, Chengwei, et al.
Veröffentlicht: (2024)
Data-efficient Meta-models for Evaluation of Context-based Questions and Answers in LLMs
von: Belikova, Julia, et al.
Veröffentlicht: (2025)
von: Belikova, Julia, et al.
Veröffentlicht: (2025)
CoDiQ: Test-Time Scaling for Controllable Difficult Question Generation
von: Peng, Zhongyuan, et al.
Veröffentlicht: (2026)
von: Peng, Zhongyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
von: Joshi, Vinay, et al.
Veröffentlicht: (2025) -
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
von: Ray, Pretam, et al.
Veröffentlicht: (2026) -
Instella: Fully Open Language Models with Stellar Performance
von: Liu, Jiang, et al.
Veröffentlicht: (2025) -
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
von: Mishra, Prakamya, et al.
Veröffentlicht: (2025) -
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)