Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mahran, Mariam, Simbeck, Katharina |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
GPT and Prejudice: A Sparse Approach to Understanding Learned Representations in Large Language Models
par: Mahran, Mariam, et autres
Publié: (2025)
par: Mahran, Mariam, et autres
Publié: (2025)
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
par: Simbeck, Katharina, et autres
Publié: (2025)
par: Simbeck, Katharina, et autres
Publié: (2025)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
par: Qin, Tian, et autres
Publié: (2025)
par: Qin, Tian, et autres
Publié: (2025)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
par: Mahdavi, Sadegh, et autres
Publié: (2025)
par: Mahdavi, Sadegh, et autres
Publié: (2025)
Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns?
par: Kao, Kuei-Chun, et autres
Publié: (2024)
par: Kao, Kuei-Chun, et autres
Publié: (2024)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
par: Opedal, Andreas, et autres
Publié: (2024)
par: Opedal, Andreas, et autres
Publié: (2024)
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models
par: Era, Jalisha Jashim, et autres
Publié: (2025)
par: Era, Jalisha Jashim, et autres
Publié: (2025)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
par: Khan, Zaid, et autres
Publié: (2025)
par: Khan, Zaid, et autres
Publié: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
par: Qu, Yuxiao, et autres
Publié: (2025)
par: Qu, Yuxiao, et autres
Publié: (2025)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
par: Ouyang, Jialin
Publié: (2025)
par: Ouyang, Jialin
Publié: (2025)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
par: Petrov, Ivo, et autres
Publié: (2025)
par: Petrov, Ivo, et autres
Publié: (2025)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
par: Chen, Nuo, et autres
Publié: (2024)
par: Chen, Nuo, et autres
Publié: (2024)
Augmenting Math Word Problems via Iterative Question Composing
par: Liu, Haoxiong, et autres
Publié: (2024)
par: Liu, Haoxiong, et autres
Publié: (2024)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
par: Chen, Zui, et autres
Publié: (2024)
par: Chen, Zui, et autres
Publié: (2024)
Deconfounded Causality-aware Parameter-Efficient Fine-Tuning for Problem-Solving Improvement of LLMs
par: Wang, Ruoyu, et autres
Publié: (2024)
par: Wang, Ruoyu, et autres
Publié: (2024)
FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline
par: Seegmiller, Parker, et autres
Publié: (2025)
par: Seegmiller, Parker, et autres
Publié: (2025)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
par: Singhi, Nishad, et autres
Publié: (2025)
par: Singhi, Nishad, et autres
Publié: (2025)
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
par: Arbabi, Alireza, et autres
Publié: (2025)
par: Arbabi, Alireza, et autres
Publié: (2025)
Do Multilingual LLMs Think In English?
par: Schut, Lisa, et autres
Publié: (2025)
par: Schut, Lisa, et autres
Publié: (2025)
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
par: Agarwal, Mehul, et autres
Publié: (2026)
par: Agarwal, Mehul, et autres
Publié: (2026)
Can Generative AI Solve Your In-Context Learning Problem? A Martingale Perspective
par: Jesson, Andrew, et autres
Publié: (2024)
par: Jesson, Andrew, et autres
Publié: (2024)
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
par: Li, Junsong, et autres
Publié: (2025)
par: Li, Junsong, et autres
Publié: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
par: Huang, Kaixuan, et autres
Publié: (2025)
par: Huang, Kaixuan, et autres
Publié: (2025)
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
par: Wang, Peiyi, et autres
Publié: (2023)
par: Wang, Peiyi, et autres
Publié: (2023)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
par: Wang, Zengzhi, et autres
Publié: (2023)
par: Wang, Zengzhi, et autres
Publié: (2023)
MegaMath: Pushing the Limits of Open Math Corpora
par: Zhou, Fan, et autres
Publié: (2025)
par: Zhou, Fan, et autres
Publié: (2025)
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
par: Wang, Xiaoxuan, et autres
Publié: (2023)
par: Wang, Xiaoxuan, et autres
Publié: (2023)
The Impact of Inference Acceleration on Bias of LLMs
par: Kirsten, Elisabeth, et autres
Publié: (2024)
par: Kirsten, Elisabeth, et autres
Publié: (2024)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
par: Pecher, Branislav, et autres
Publié: (2026)
par: Pecher, Branislav, et autres
Publié: (2026)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
par: Hao, Yuren, et autres
Publié: (2025)
par: Hao, Yuren, et autres
Publié: (2025)
Debiasing Multilingual LLMs in Cross-lingual Latent Space
par: Peng, Qiwei, et autres
Publié: (2025)
par: Peng, Qiwei, et autres
Publié: (2025)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
par: Yoo, Haneul, et autres
Publié: (2024)
par: Yoo, Haneul, et autres
Publié: (2024)
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
par: Li, Chengpeng, et autres
Publié: (2023)
par: Li, Chengpeng, et autres
Publié: (2023)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
par: Toshniwal, Shubham, et autres
Publié: (2024)
par: Toshniwal, Shubham, et autres
Publié: (2024)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
par: Alam, Firoj, et autres
Publié: (2026)
par: Alam, Firoj, et autres
Publié: (2026)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
par: Saha, Swarnadeep, et autres
Publié: (2023)
par: Saha, Swarnadeep, et autres
Publié: (2023)
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity
par: Burnat, Florian A. D., et autres
Publié: (2026)
par: Burnat, Florian A. D., et autres
Publié: (2026)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
par: Liu, Zihan, et autres
Publié: (2024)
par: Liu, Zihan, et autres
Publié: (2024)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
par: Mahabadi, Rabeeh Karimi, et autres
Publié: (2025)
par: Mahabadi, Rabeeh Karimi, et autres
Publié: (2025)
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
par: Albalak, Alon, et autres
Publié: (2025)
par: Albalak, Alon, et autres
Publié: (2025)
Documents similaires
-
GPT and Prejudice: A Sparse Approach to Understanding Learned Representations in Large Language Models
par: Mahran, Mariam, et autres
Publié: (2025) -
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
par: Simbeck, Katharina, et autres
Publié: (2025) -
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
par: Qin, Tian, et autres
Publié: (2025) -
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
par: Mahdavi, Sadegh, et autres
Publié: (2025) -
Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns?
par: Kao, Kuei-Chun, et autres
Publié: (2024)