Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange
Fuente:
arXiv
Saved in:
| Main Authors: | Satpute, Ankit, Giessing, Noah, Greiner-Petter, Andre, Schubotz, Moritz, Teschke, Olaf, Aizawa, Akiko, Gipp, Bela |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aspect-Aware Content-Based Recommendations for Mathematical Research Papers
by: Satpute, Ankit, et al.
Published: (2026)
by: Satpute, Ankit, et al.
Published: (2026)
Taxonomy of Mathematical Plagiarism
by: Satpute, Ankit, et al.
Published: (2024)
by: Satpute, Ankit, et al.
Published: (2024)
Reducing the climate impact of data portals: a case study
by: Gießing, Noah, et al.
Published: (2024)
by: Gießing, Noah, et al.
Published: (2024)
LLM-supported document separation for printed reviews from zbMATH Open
by: Pluzhnikov, Ivan, et al.
Published: (2026)
by: Pluzhnikov, Ivan, et al.
Published: (2026)
Overview of the Plagiarism Detection Task at PAN 2025
by: Greiner-Petter, André, et al.
Published: (2025)
by: Greiner-Petter, André, et al.
Published: (2025)
Making Presentation Math Computable
by: Greiner-Petter, André
Published: (2023)
by: Greiner-Petter, André
Published: (2023)
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
by: Horych, Tomas, et al.
Published: (2024)
by: Horych, Tomas, et al.
Published: (2024)
An Overview of zbMATH Open Digital Library
by: Deb, Madhurima, et al.
Published: (2024)
by: Deb, Madhurima, et al.
Published: (2024)
Beyond Chains: Bridging Large Language Models and Knowledge Bases in Complex Question Answering
by: Zhu, Yihua, et al.
Published: (2025)
by: Zhu, Yihua, et al.
Published: (2025)
Contrastive Learning Using Graph Embeddings for Domain Adaptation of Language Models in the Process Industry
by: Zhukova, Anastasia, et al.
Published: (2025)
by: Zhukova, Anastasia, et al.
Published: (2025)
From AutoRecSys to AutoRecLab: A Call to Build, Evaluate, and Govern Autonomous Recommender-Systems Research Labs
by: Beel, Joeran, et al.
Published: (2025)
by: Beel, Joeran, et al.
Published: (2025)
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
by: Boudin, Florian, et al.
Published: (2025)
by: Boudin, Florian, et al.
Published: (2025)
Challenging Tribal Knowledge -- Large Scale Measurement Campaign on Decentralized NAT Traversal
by: Trautwein, Dennis, et al.
Published: (2025)
by: Trautwein, Dennis, et al.
Published: (2025)
Link Prediction for Event Logs in the Process Industry
by: Zhukova, Anastasia, et al.
Published: (2025)
by: Zhukova, Anastasia, et al.
Published: (2025)
A Study of PHOC Spatial Region Configurations for Math Formula Retrieval
by: Langsenkamp, Matt, et al.
Published: (2024)
by: Langsenkamp, Matt, et al.
Published: (2024)
TWOLAR: a TWO-step LLM-Augmented distillation method for passage Reranking
by: Baldelli, Davide, et al.
Published: (2024)
by: Baldelli, Davide, et al.
Published: (2024)
Revisiting Bi-Encoder Neural Search: An Encoding--Searching Separation Perspective
by: Tran, Hung-Nghiep, et al.
Published: (2024)
by: Tran, Hung-Nghiep, et al.
Published: (2024)
zbMATH Open: API Solutions and Research Challenges
by: Petrera, Matteo, et al.
Published: (2021)
by: Petrera, Matteo, et al.
Published: (2021)
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
by: Huang, Youcheng, et al.
Published: (2025)
by: Huang, Youcheng, et al.
Published: (2025)
Self-Compositional Data Augmentation for Scientific Keyphrase Generation
by: Houbre, Mael, et al.
Published: (2024)
by: Houbre, Mael, et al.
Published: (2024)
Med-CoDE: Medical Critique based Disagreement Evaluation Framework
by: Gupta, Mohit, et al.
Published: (2025)
by: Gupta, Mohit, et al.
Published: (2025)
MAGPIE: Multi-Task Media-Bias Analysis Generalization for Pre-Trained Identification of Expressions
by: Horych, Tomáš, et al.
Published: (2024)
by: Horych, Tomáš, et al.
Published: (2024)
Large-Scale Measurement of NAT Traversal for the Decentralized Web: A Case Study of DCUtR in IPFS
by: Trautwein, Dennis, et al.
Published: (2026)
by: Trautwein, Dennis, et al.
Published: (2026)
Author Intent: Eliminating Ambiguity in MathML
by: Carlisle, David, et al.
Published: (2024)
by: Carlisle, David, et al.
Published: (2024)
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
by: Teixeira, Tiago, et al.
Published: (2026)
by: Teixeira, Tiago, et al.
Published: (2026)
WikiTexVC: MediaWiki's native LaTeX to MathML converter for Wikipedia
by: Stegmüller, Johannes, et al.
Published: (2024)
by: Stegmüller, Johannes, et al.
Published: (2024)
Large Search Model: Redefining Search Stack in the Era of LLMs
by: Wang, Liang, et al.
Published: (2023)
by: Wang, Liang, et al.
Published: (2023)
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
by: Alshammari, Shaden, et al.
Published: (2026)
by: Alshammari, Shaden, et al.
Published: (2026)
Enhancing Math Learning in an LMS Using AI-Driven Question Recommendations
by: Råmunddal, Justus
Published: (2025)
by: Råmunddal, Justus
Published: (2025)
No Stupid Questions: An Analysis of Question Query Generation for Citation Recommendation
by: Zimmerman, Brian D., et al.
Published: (2025)
by: Zimmerman, Brian D., et al.
Published: (2025)
Comparing RAG and GraphRAG for Page-Level Retrieval Question Answering on Math Textbook
by: Chen, Eason, et al.
Published: (2025)
by: Chen, Eason, et al.
Published: (2025)
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
by: Guan, Xinyu, et al.
Published: (2025)
by: Guan, Xinyu, et al.
Published: (2025)
Engaging Students Through Math Competitions
by: Bajnok, Bela
Published: (2024)
by: Bajnok, Bela
Published: (2024)
SBI-RAG: Enhancing Math Word Problem Solving for Students through Schema-Based Instruction and Retrieval-Augmented Generation
by: Dixit, Prakhar, et al.
Published: (2024)
by: Dixit, Prakhar, et al.
Published: (2024)
Can Large Language Models Assess Serendipity in Recommender Systems?
by: Tokutake, Yu, et al.
Published: (2024)
by: Tokutake, Yu, et al.
Published: (2024)
Full-Stack Optimized Large Language Models for Lifelong Sequential Behavior Comprehension in Recommendation
by: Shan, Rong, et al.
Published: (2025)
by: Shan, Rong, et al.
Published: (2025)
A Reproducibility and Generalizability Study of Large Language Models for Query Generation
by: Staudinger, Moritz, et al.
Published: (2024)
by: Staudinger, Moritz, et al.
Published: (2024)
LLMs Can Patch Up Missing Relevance Judgments in Evaluation
by: Upadhyay, Shivani, et al.
Published: (2024)
by: Upadhyay, Shivani, et al.
Published: (2024)
Can Large Language Models Detect Rumors on Social Media?
by: Liu, Qiang, et al.
Published: (2024)
by: Liu, Qiang, et al.
Published: (2024)
Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
by: Kumar, Sunisth, et al.
Published: (2026)
by: Kumar, Sunisth, et al.
Published: (2026)
Similar Items
-
Aspect-Aware Content-Based Recommendations for Mathematical Research Papers
by: Satpute, Ankit, et al.
Published: (2026) -
Taxonomy of Mathematical Plagiarism
by: Satpute, Ankit, et al.
Published: (2024) -
Reducing the climate impact of data portals: a case study
by: Gießing, Noah, et al.
Published: (2024) -
LLM-supported document separation for printed reviews from zbMATH Open
by: Pluzhnikov, Ivan, et al.
Published: (2026) -
Overview of the Plagiarism Detection Task at PAN 2025
by: Greiner-Petter, André, et al.
Published: (2025)