Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Al-Khalili, Zena, Howell, Nick, Klakow, Dietrich |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Human Speech Perception in Noise: Can Large Language Models Paraphrase to Improve It?
par: Chingacham, Anupama, et autres
Publié: (2024)
par: Chingacham, Anupama, et autres
Publié: (2024)
Exploring the Effectiveness and Consistency of Task Selection in Intermediate-Task Transfer Learning
par: Lin, Pin-Jie, et autres
Publié: (2024)
par: Lin, Pin-Jie, et autres
Publié: (2024)
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
par: Yu, Dian, et autres
Publié: (2024)
par: Yu, Dian, et autres
Publié: (2024)
Improving Semantic Understanding in Speech Language Models via Brain-tuning
par: Moussa, Omer, et autres
Publié: (2024)
par: Moussa, Omer, et autres
Publié: (2024)
Utilizing Multimodal Data for Edge Case Robust Call-sign Recognition and Understanding
par: Blatt, Alexander, et autres
Publié: (2024)
par: Blatt, Alexander, et autres
Publié: (2024)
A Preference-driven Paradigm for Enhanced Translation with Large Language Models
par: Zhu, Dawei, et autres
Publié: (2024)
par: Zhu, Dawei, et autres
Publié: (2024)
Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?
par: Zhu, Dawei, et autres
Publié: (2024)
par: Zhu, Dawei, et autres
Publié: (2024)
IGC: Integrating a Gated Calculator into an LLM to Solve Arithmetic Tasks Reliably and Efficiently
par: Dietz, Florian, et autres
Publié: (2025)
par: Dietz, Florian, et autres
Publié: (2025)
Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding
par: Belay, Tadesse Destaw, et autres
Publié: (2024)
par: Belay, Tadesse Destaw, et autres
Publié: (2024)
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
par: Gautam, Vagrant, et autres
Publié: (2024)
par: Gautam, Vagrant, et autres
Publié: (2024)
Gap-Filling Prompting Enhances Code-Assisted Mathematical Reasoning
par: Mohammadkhani, Mohammad Ghiasvand
Publié: (2024)
par: Mohammadkhani, Mohammad Ghiasvand
Publié: (2024)
CodeMind: Evaluating Large Language Models for Code Reasoning
par: Liu, Changshu, et autres
Publié: (2024)
par: Liu, Changshu, et autres
Publié: (2024)
On the Encoding of Gender in Transformer-based ASR Representations
par: Krishnan, Aravind, et autres
Publié: (2024)
par: Krishnan, Aravind, et autres
Publié: (2024)
The Hidden Space of Transformer Language Adapters
par: Alabi, Jesujoba O., et autres
Publié: (2024)
par: Alabi, Jesujoba O., et autres
Publié: (2024)
Large Language Models for Mathematical Reasoning: Progresses and Challenges
par: Ahn, Janice, et autres
Publié: (2024)
par: Ahn, Janice, et autres
Publié: (2024)
Aligned Probing: Relating Toxic Behavior and Model Internals
par: Waldis, Andreas, et autres
Publié: (2025)
par: Waldis, Andreas, et autres
Publié: (2025)
PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing Tasks
par: Liu, Yunuo, et autres
Publié: (2025)
par: Liu, Yunuo, et autres
Publié: (2025)
EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation
par: Tonja, Atnafu Lambebo, et autres
Publié: (2024)
par: Tonja, Atnafu Lambebo, et autres
Publié: (2024)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
par: Amjad, Husnain, et autres
Publié: (2026)
par: Amjad, Husnain, et autres
Publié: (2026)
Understanding "Democratization" in NLP and ML Research
par: Subramonian, Arjun, et autres
Publié: (2024)
par: Subramonian, Arjun, et autres
Publié: (2024)
Joint vs Sequential Speaker-Role Detection and Automatic Speech Recognition for Air-traffic Control
par: Blatt, Alexander, et autres
Publié: (2024)
par: Blatt, Alexander, et autres
Publié: (2024)
MathRobust-LV: Evaluation of Large Language Models' Robustness to Linguistic Variations in Mathematical Reasoning
par: Kirtane, Neeraja, et autres
Publié: (2025)
par: Kirtane, Neeraja, et autres
Publié: (2025)
Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering
par: Ferguson, Nick, et autres
Publié: (2025)
par: Ferguson, Nick, et autres
Publié: (2025)
Dual Instruction Tuning with Large Language Models for Mathematical Reasoning
par: Zhou, Yongwei, et autres
Publié: (2024)
par: Zhou, Yongwei, et autres
Publié: (2024)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
par: Subramonian, Arjun, et autres
Publié: (2025)
par: Subramonian, Arjun, et autres
Publié: (2025)
A Survey on Large Language Models for Mathematical Reasoning
par: Wang, Peng-Yuan, et autres
Publié: (2025)
par: Wang, Peng-Yuan, et autres
Publié: (2025)
Saar-Voice: A Multi-Speaker Saarbrücken Dialect Speech Corpus
par: Oberkircher, Lena S., et autres
Publié: (2026)
par: Oberkircher, Lena S., et autres
Publié: (2026)
AAdaM at SemEval-2024 Task 1: Augmentation and Adaptation for Multilingual Semantic Textual Relatedness
par: Zhang, Miaoran, et autres
Publié: (2024)
par: Zhang, Miaoran, et autres
Publié: (2024)
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
par: Xue, Boyang, et autres
Publié: (2025)
par: Xue, Boyang, et autres
Publié: (2025)
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
par: Das, Debrup, et autres
Publié: (2024)
par: Das, Debrup, et autres
Publié: (2024)
Bridging the Culture Gap: A Framework for LLM-Driven Socio-Cultural Localization of Math Word Problems in Low-Resource Languages
par: Azime, Israel Abebe, et autres
Publié: (2025)
par: Azime, Israel Abebe, et autres
Publié: (2025)
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning
par: Yang, Rem, et autres
Publié: (2025)
par: Yang, Rem, et autres
Publié: (2025)
Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
par: García-de-Herreros, Paloma, et autres
Publié: (2025)
par: García-de-Herreros, Paloma, et autres
Publié: (2025)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
par: Zhao, Jun, et autres
Publié: (2024)
par: Zhao, Jun, et autres
Publié: (2024)
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
par: Lin, Honglin, et autres
Publié: (2025)
par: Lin, Honglin, et autres
Publié: (2025)
SMRC: Aligning Large Language Models with Student Reasoning for Mathematical Error Correction
par: Zeng, Biaojie, et autres
Publié: (2025)
par: Zeng, Biaojie, et autres
Publié: (2025)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
par: Shi, Wenhao, et autres
Publié: (2024)
par: Shi, Wenhao, et autres
Publié: (2024)
Can Language Models Rival Mathematics Students? Evaluating Mathematical Reasoning through Textual Manipulation and Human Experiments
par: Nikolaiev, Andrii, et autres
Publié: (2024)
par: Nikolaiev, Andrii, et autres
Publié: (2024)
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
par: Li, Xiaoyuan, et autres
Publié: (2024)
par: Li, Xiaoyuan, et autres
Publié: (2024)
MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning
par: Yin, Shuo, et autres
Publié: (2024)
par: Yin, Shuo, et autres
Publié: (2024)
Documents similaires
-
Human Speech Perception in Noise: Can Large Language Models Paraphrase to Improve It?
par: Chingacham, Anupama, et autres
Publié: (2024) -
Exploring the Effectiveness and Consistency of Task Selection in Intermediate-Task Transfer Learning
par: Lin, Pin-Jie, et autres
Publié: (2024) -
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
par: Yu, Dian, et autres
Publié: (2024) -
Improving Semantic Understanding in Speech Language Models via Brain-tuning
par: Moussa, Omer, et autres
Publié: (2024) -
Utilizing Multimodal Data for Edge Case Robust Call-sign Recognition and Understanding
par: Blatt, Alexander, et autres
Publié: (2024)