PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yiming, Zhang, Pei, Tang, Jialong, Wei, Haoran, Yang, Baosong, Wang, Rui, Sun, Chenshu, Sun, Feitong, Zhang, Jiran, Wu, Junxuan, Cang, Qiqian, Zhang, Yichang, Huang, Fei, Lin, Junyang, Zhou, Jingren |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
by: Zhang, Yidan, et al.
Published: (2024)
by: Zhang, Yidan, et al.
Published: (2024)
Direct Simultaneous Translation Activation for Large Audio-Language Models
by: Zhang, Pei, et al.
Published: (2025)
by: Zhang, Pei, et al.
Published: (2025)
Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
by: Wang, Yiming, et al.
Published: (2025)
by: Wang, Yiming, et al.
Published: (2025)
ConText: Driving In-context Learning for Text Removal and Segmentation
by: Zhang, Fei, et al.
Published: (2025)
by: Zhang, Fei, et al.
Published: (2025)
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
by: Liu, Yikang, et al.
Published: (2025)
by: Liu, Yikang, et al.
Published: (2025)
CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
by: Zhang, Pei, et al.
Published: (2025)
by: Zhang, Pei, et al.
Published: (2025)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
Meta-Reasoning: Semantics-Symbol Deconstruction for Large Language Models
by: Wang, Yiming, et al.
Published: (2023)
by: Wang, Yiming, et al.
Published: (2023)
Towards Cross-lingual Values Judgment: A Consensus-Pluralism Perspective
by: Chen, Yukun, et al.
Published: (2026)
by: Chen, Yukun, et al.
Published: (2026)
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
by: Zhang, Yanzhao, et al.
Published: (2025)
by: Zhang, Yanzhao, et al.
Published: (2025)
Enhancing LLM Language Adaption through Cross-lingual In-Context Pre-training
by: Wu, Linjuan, et al.
Published: (2025)
by: Wu, Linjuan, et al.
Published: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing
by: Xiang, Hao, et al.
Published: (2025)
by: Xiang, Hao, et al.
Published: (2025)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
by: Fang, Meng, et al.
Published: (2024)
by: Fang, Meng, et al.
Published: (2024)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
by: Li, Can, et al.
Published: (2025)
by: Li, Can, et al.
Published: (2025)
Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation
by: Zhang, Collin, et al.
Published: (2025)
by: Zhang, Collin, et al.
Published: (2025)
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
by: Yang, An, et al.
Published: (2024)
by: Yang, An, et al.
Published: (2024)
Qwen3-ASR Technical Report
by: Shi, Xian, et al.
Published: (2026)
by: Shi, Xian, et al.
Published: (2026)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
Not All Languages are Equal: Insights into Multilingual Retrieval-Augmented Generation
by: Wu, Suhang, et al.
Published: (2024)
by: Wu, Suhang, et al.
Published: (2024)
Re$^2$Math: Benchmarking Theorem Retrieval in Research-Level Mathematics
by: Lyu, Zicheng, et al.
Published: (2026)
by: Lyu, Zicheng, et al.
Published: (2026)
Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning
by: Zhuang, Wenwen, et al.
Published: (2024)
by: Zhuang, Wenwen, et al.
Published: (2024)
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator
by: Fan, Zhihao, et al.
Published: (2024)
by: Fan, Zhihao, et al.
Published: (2024)
Retrieved In-Context Principles from Previous Mistakes
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
by: Zhang, Zhenru, et al.
Published: (2025)
by: Zhang, Zhenru, et al.
Published: (2025)
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
by: Qiao, Runqi, et al.
Published: (2025)
by: Qiao, Runqi, et al.
Published: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
by: Tang, Zhengyang, et al.
Published: (2024)
by: Tang, Zhengyang, et al.
Published: (2024)
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
by: Xue, Boyang, et al.
Published: (2025)
by: Xue, Boyang, et al.
Published: (2025)
SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
by: Deng, Boyi, et al.
Published: (2025)
by: Deng, Boyi, et al.
Published: (2025)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
by: Wang, Junyang, et al.
Published: (2025)
by: Wang, Junyang, et al.
Published: (2025)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
by: Wang, Junyang, et al.
Published: (2025)
by: Wang, Junyang, et al.
Published: (2025)
Unlocking Interpretability for RF Sensing: A Complex-Valued White-Box Transformer
by: Zhang, Xie, et al.
Published: (2025)
by: Zhang, Xie, et al.
Published: (2025)
A Probabilistic Inference Scaling Theory for LLM Self-Correction
by: Yang, Zhe, et al.
Published: (2025)
by: Yang, Zhe, et al.
Published: (2025)
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
by: Yang, Zhe, et al.
Published: (2024)
by: Yang, Zhe, et al.
Published: (2024)
Similar Items
-
P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
by: Zhang, Yidan, et al.
Published: (2024) -
Direct Simultaneous Translation Activation for Large Audio-Language Models
by: Zhang, Pei, et al.
Published: (2025) -
Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding
by: Wang, Yiming, et al.
Published: (2025) -
ConText: Driving In-context Learning for Text Removal and Segmentation
by: Zhang, Fei, et al.
Published: (2025) -
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
by: Liu, Yikang, et al.
Published: (2025)