IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
Fuente:
arXiv
Saved in:
| Main Authors: | Biyani, Param, Kirtania, Shashank, Bajpai, Yasharth, Gulwani, Sumit, Tiwari, Ashish |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Language Agents through BREW
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
Exploring Interaction Patterns for Debugging: Enhancing Conversational Capabilities of AI-assistants
by: Chopra, Bhavya, et al.
Published: (2024)
by: Chopra, Bhavya, et al.
Published: (2024)
TableTalk: Scaffolding Spreadsheet Development with a Language Agent
by: Liang, Jenny T., et al.
Published: (2025)
by: Liang, Jenny T., et al.
Published: (2025)
MetaReflection: Learning Instructions for Language Agents using Past Reflections
by: Gupta, Priyanshu, et al.
Published: (2024)
by: Gupta, Priyanshu, et al.
Published: (2024)
TEN: Table Explicitization, Neurosymbolically
by: Mehrotra, Nikita, et al.
Published: (2025)
by: Mehrotra, Nikita, et al.
Published: (2025)
Steering LLMs for Formal Theorem Proving
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
by: Mhatre, Sanket, et al.
Published: (2025)
by: Mhatre, Sanket, et al.
Published: (2025)
Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
MathAtlas: A Benchmark for Autoformalization in the Wild
by: Patel, Nilay, et al.
Published: (2026)
by: Patel, Nilay, et al.
Published: (2026)
LOGIC-LM++: Multi-Step Refinement for Symbolic Formulations
by: Kirtania, Shashank, et al.
Published: (2024)
by: Kirtania, Shashank, et al.
Published: (2024)
STACKFEED: Structured Textual Actor-Critic Knowledge Base Editing with FeedBack
by: Kirtania, Shashank, et al.
Published: (2024)
by: Kirtania, Shashank, et al.
Published: (2024)
Why AI Agents Still Need You: Findings from Developer-Agent Collaborations in the Wild
by: Kumar, Aayush, et al.
Published: (2025)
by: Kumar, Aayush, et al.
Published: (2025)
Enhancing Creativity in Large Language Models through Associative Thinking Strategies
by: Mehrotra, Pronita, et al.
Published: (2024)
by: Mehrotra, Pronita, et al.
Published: (2024)
ConDABench: Interactive Evaluation of Language Models for Data Analysis
by: Dutta, Avik, et al.
Published: (2025)
by: Dutta, Avik, et al.
Published: (2025)
RoMath: A Mathematical Reasoning Benchmark in Romanian
by: Cosma, Adrian, et al.
Published: (2024)
by: Cosma, Adrian, et al.
Published: (2024)
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
by: Glazer, Elliot, et al.
Published: (2024)
by: Glazer, Elliot, et al.
Published: (2024)
Diffusion is a code repair operator and generator
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Autoformalizer with Tool Feedback
by: Guo, Qi, et al.
Published: (2025)
by: Guo, Qi, et al.
Published: (2025)
Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
An Empirical Investigation of Robustness in Large Language Models under Tabular Distortions
by: Dutta, Avik, et al.
Published: (2026)
by: Dutta, Avik, et al.
Published: (2026)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions
by: Liang, Zhenwen, et al.
Published: (2024)
by: Liang, Zhenwen, et al.
Published: (2024)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
by: Tang, Zhengyang, et al.
Published: (2024)
by: Tang, Zhengyang, et al.
Published: (2024)
Do Code Models Suffer from the Dunning-Kruger Effect?
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
by: Guo, Zixian, et al.
Published: (2025)
by: Guo, Zixian, et al.
Published: (2025)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
by: Tong, Yuxuan, et al.
Published: (2024)
by: Tong, Yuxuan, et al.
Published: (2024)
Towards a Common Framework for Autoformalization
by: Mensfelt, Agnieszka, et al.
Published: (2025)
by: Mensfelt, Agnieszka, et al.
Published: (2025)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
by: Fang, Meng, et al.
Published: (2024)
by: Fang, Meng, et al.
Published: (2024)
MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models
by: Peng, Shuai, et al.
Published: (2024)
by: Peng, Shuai, et al.
Published: (2024)
EvolMathEval: Towards Evolvable Benchmarks for Mathematical Reasoning via Evolutionary Testing
by: Wang, Shengbo, et al.
Published: (2025)
by: Wang, Shengbo, et al.
Published: (2025)
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
by: Li, Can, et al.
Published: (2025)
by: Li, Can, et al.
Published: (2025)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
by: Pei, Qizhi, et al.
Published: (2025)
by: Pei, Qizhi, et al.
Published: (2025)
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task
by: Yan, Yuchen, et al.
Published: (2025)
by: Yan, Yuchen, et al.
Published: (2025)
Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
by: Du, Wei, et al.
Published: (2025)
by: Du, Wei, et al.
Published: (2025)
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
by: Shao, Zhihong, et al.
Published: (2025)
by: Shao, Zhihong, et al.
Published: (2025)
MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis
by: Yu, Zixiong, et al.
Published: (2026)
by: Yu, Zixiong, et al.
Published: (2026)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
Tabularis Formatus: Predictive Formatting for Tables
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Similar Items
-
Improving Language Agents through BREW
by: Kirtania, Shashank, et al.
Published: (2025) -
Exploring Interaction Patterns for Debugging: Enhancing Conversational Capabilities of AI-assistants
by: Chopra, Bhavya, et al.
Published: (2024) -
TableTalk: Scaffolding Spreadsheet Development with a Language Agent
by: Liang, Jenny T., et al.
Published: (2025) -
MetaReflection: Learning Instructions for Language Agents using Past Reflections
by: Gupta, Priyanshu, et al.
Published: (2024) -
TEN: Table Explicitization, Neurosymbolically
by: Mehrotra, Nikita, et al.
Published: (2025)