Saved in:
| Main Authors: | Rastogi, Ritvik, Dharashivkar, Sachin, Varma, Sandeep |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.08665 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aryabhata 2: Scaling Reinforcement Learning for Advanced STEM Reasoning
by: Rastogi, Ritvik, et al.
Published: (2026)
by: Rastogi, Ritvik, et al.
Published: (2026)
MathDivide: Improved mathematical reasoning by large language models
by: Srivastava, Saksham Sahai, et al.
Published: (2024)
by: Srivastava, Saksham Sahai, et al.
Published: (2024)
The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories
by: Shah, Raj Sanjay, et al.
Published: (2025)
by: Shah, Raj Sanjay, et al.
Published: (2025)
A Novel Approach to Image EEG Sleep Data for Improving Quality of Life in Patients Suffering From Brain Injuries Using DreamDiffusion
by: Fahim, David, et al.
Published: (2024)
by: Fahim, David, et al.
Published: (2024)
fact check AI at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-checked Claim Retrieval
by: Rastogi, Pranshu
Published: (2025)
by: Rastogi, Pranshu
Published: (2025)
PROVEX: Enhancing SOC Analyst Trust with Explainable Provenance-Based IDS
by: Dhanuka, Devang, et al.
Published: (2025)
by: Dhanuka, Devang, et al.
Published: (2025)
DEI: Diversity in Evolutionary Inference for Quality-Diversity Search
by: Donaghy, John, et al.
Published: (2026)
by: Donaghy, John, et al.
Published: (2026)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
by: Glazer, Elliot, et al.
Published: (2024)
by: Glazer, Elliot, et al.
Published: (2024)
Domain Generalization In Robust Invariant Representation
by: Gupta, Gauri, et al.
Published: (2023)
by: Gupta, Gauri, et al.
Published: (2023)
ContraPrompt: Contrastive Prompt Optimization via Dyadic Reasoning Trace Analysis
by: Rishav, Rishav, et al.
Published: (2026)
by: Rishav, Rishav, et al.
Published: (2026)
CoDream: Exchanging dreams instead of models for federated aggregation with heterogeneous models
by: Singh, Abhishek, et al.
Published: (2024)
by: Singh, Abhishek, et al.
Published: (2024)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
by: Mitra, Arindam, et al.
Published: (2024)
by: Mitra, Arindam, et al.
Published: (2024)
Chatsparent: An Interactive System for Detecting and Mitigating Cognitive Fatigue in LLMs
by: Marwah, Riju, et al.
Published: (2025)
by: Marwah, Riju, et al.
Published: (2025)
TimeSeriesExam: A time series understanding exam
by: Cai, Yifu, et al.
Published: (2024)
by: Cai, Yifu, et al.
Published: (2024)
MegaMath: Pushing the Limits of Open Math Corpora
by: Zhou, Fan, et al.
Published: (2025)
by: Zhou, Fan, et al.
Published: (2025)
Dissociating language and thought in large language models
by: Mahowald, Kyle, et al.
Published: (2023)
by: Mahowald, Kyle, et al.
Published: (2023)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
by: Tian, Shi-Yu, et al.
Published: (2025)
by: Tian, Shi-Yu, et al.
Published: (2025)
DOoM: Difficult Olympiads of Math
by: Kuleshov, Ilya, et al.
Published: (2025)
by: Kuleshov, Ilya, et al.
Published: (2025)
GamiBench: Evaluating Spatial Reasoning and 2D-to-3D Planning Capabilities of MLLMs with Origami Folding Tasks
by: Spencer, Ryan, et al.
Published: (2025)
by: Spencer, Ryan, et al.
Published: (2025)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
by: Liu, Xianyang, et al.
Published: (2025)
by: Liu, Xianyang, et al.
Published: (2025)
MathMistake Checker: A Comprehensive Demonstration for Step-by-Step Math Problem Mistake Finding by Prompt-Guided LLMs
by: Zhang, Tianyang, et al.
Published: (2025)
by: Zhang, Tianyang, et al.
Published: (2025)
The wall confronting large language models
by: Coveney, Peter V., et al.
Published: (2025)
by: Coveney, Peter V., et al.
Published: (2025)
Pessimistic Verification for Open Ended Math Questions
by: Huang, Yanxing, et al.
Published: (2025)
by: Huang, Yanxing, et al.
Published: (2025)
Neuro-Symbolic Data Generation for Math Reasoning
by: Li, Zenan, et al.
Published: (2024)
by: Li, Zenan, et al.
Published: (2024)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward Passes
by: Christ, Bryan R., et al.
Published: (2024)
by: Christ, Bryan R., et al.
Published: (2024)
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
by: Li, Chengpeng, et al.
Published: (2023)
by: Li, Chengpeng, et al.
Published: (2023)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
by: Chen, Nuo, et al.
Published: (2024)
by: Chen, Nuo, et al.
Published: (2024)
Recovering implicit physics model under real-world constraints
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
MathBuddy: A Multimodal System for Affective Math Tutoring
by: Kar, Debanjana, et al.
Published: (2025)
by: Kar, Debanjana, et al.
Published: (2025)
The effect of fine-tuning on language model toxicity
by: Hawkins, Will, et al.
Published: (2024)
by: Hawkins, Will, et al.
Published: (2024)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)
by: Van Long, Phuoc Pham, et al.
Published: (2023)
Fluent dreaming for language models
by: Thompson, T. Ben, et al.
Published: (2024)
by: Thompson, T. Ben, et al.
Published: (2024)
Algorithmic progress in language models
by: Ho, Anson, et al.
Published: (2024)
by: Ho, Anson, et al.
Published: (2024)
Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation
by: Borah, Arnabh, et al.
Published: (2025)
by: Borah, Arnabh, et al.
Published: (2025)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
by: Liu, Zihan, et al.
Published: (2024)
by: Liu, Zihan, et al.
Published: (2024)
MARGE: Improving Math Reasoning for LLMs with Guided Exploration
by: Gao, Jingyue, et al.
Published: (2025)
by: Gao, Jingyue, et al.
Published: (2025)
MathConstruct: Challenging LLM Reasoning with Constructive Proofs
by: Balunović, Mislav, et al.
Published: (2025)
by: Balunović, Mislav, et al.
Published: (2025)
Similar Items
-
Aryabhata 2: Scaling Reinforcement Learning for Advanced STEM Reasoning
by: Rastogi, Ritvik, et al.
Published: (2026) -
MathDivide: Improved mathematical reasoning by large language models
by: Srivastava, Saksham Sahai, et al.
Published: (2024) -
The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories
by: Shah, Raj Sanjay, et al.
Published: (2025) -
A Novel Approach to Image EEG Sleep Data for Improving Quality of Life in Patients Suffering From Brain Injuries Using DreamDiffusion
by: Fahim, David, et al.
Published: (2024) -
fact check AI at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-checked Claim Retrieval
by: Rastogi, Pranshu
Published: (2025)