VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Jingkun, Zhan, Runzhe, Li, Yang, Sun, Di, Chan, Hou Pong, Chao, Lidia S., Wong, Derek F. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model
von: Xu, Haoyun, et al.
Veröffentlicht: (2024)
von: Xu, Haoyun, et al.
Veröffentlicht: (2024)
Intrinsic Model Weaknesses: How Priming Attacks Unveil Vulnerabilities in Large Language Models
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
von: Huang, Yuyi, et al.
Veröffentlicht: (2025)
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
Prefix Text as a Yarn: Eliciting Non-English Alignment in Foundation Language Model
von: Zhan, Runzhe, et al.
Veröffentlicht: (2024)
von: Zhan, Runzhe, et al.
Veröffentlicht: (2024)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
Rethinking Prompt-based Debiasing in Large Language Models
von: Yang, Xinyi, et al.
Veröffentlicht: (2025)
von: Yang, Xinyi, et al.
Veröffentlicht: (2025)
A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
von: Wu, Junchao, et al.
Veröffentlicht: (2023)
von: Wu, Junchao, et al.
Veröffentlicht: (2023)
Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models
von: Peng, Shuai, et al.
Veröffentlicht: (2024)
von: Peng, Shuai, et al.
Veröffentlicht: (2024)
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
von: Chen, Xin, et al.
Veröffentlicht: (2026)
von: Chen, Xin, et al.
Veröffentlicht: (2026)
Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
von: Chen, Guanhua, et al.
Veröffentlicht: (2026)
von: Chen, Guanhua, et al.
Veröffentlicht: (2026)
RoMath: A Mathematical Reasoning Benchmark in Romanian
von: Cosma, Adrian, et al.
Veröffentlicht: (2024)
von: Cosma, Adrian, et al.
Veröffentlicht: (2024)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation
von: Sun, Yanming, et al.
Veröffentlicht: (2025)
von: Sun, Yanming, et al.
Veröffentlicht: (2025)
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
von: Chen, Xin, et al.
Veröffentlicht: (2025)
von: Chen, Xin, et al.
Veröffentlicht: (2025)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
von: Wei, Chengwei, et al.
Veröffentlicht: (2025)
von: Wei, Chengwei, et al.
Veröffentlicht: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
von: Sobhani, Mahbub E, et al.
Veröffentlicht: (2025)
von: Sobhani, Mahbub E, et al.
Veröffentlicht: (2025)
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation
von: Wei, Hu, et al.
Veröffentlicht: (2025)
von: Wei, Hu, et al.
Veröffentlicht: (2025)
JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2023)
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2023)
EternalMath: A Living Benchmark of Frontier Mathematics that Evolves with Human Discovery
von: Ma, Jicheng, et al.
Veröffentlicht: (2026)
von: Ma, Jicheng, et al.
Veröffentlicht: (2026)
ExGRPO: Learning to Reason from Experience
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025)
A Two-Stage Prediction-Aware Contrastive Learning Framework for Multi-Intent NLU
von: Chen, Guanhua, et al.
Veröffentlicht: (2024)
von: Chen, Guanhua, et al.
Veröffentlicht: (2024)
ManiTweet: A New Benchmark for Identifying Manipulation of News on Social Media
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2023)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2023)
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2025)
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2025)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
von: Lu, Pan, et al.
Veröffentlicht: (2023)
von: Lu, Pan, et al.
Veröffentlicht: (2023)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
AMERICANO: Argument Generation with Discourse-driven Decomposition and Agent Interaction
von: Hu, Zhe, et al.
Veröffentlicht: (2023)
von: Hu, Zhe, et al.
Veröffentlicht: (2023)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
von: Wang, Shanshan, et al.
Veröffentlicht: (2025)
Towards an AI Musician: Synthesizing Sheet Music Problems for Musical Reasoning
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
von: Huang, Yuyi, et al.
Veröffentlicht: (2025) -
Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model
von: Xu, Haoyun, et al.
Veröffentlicht: (2024) -
Intrinsic Model Weaknesses: How Priming Attacks Unveil Vulnerabilities in Large Language Models
von: Huang, Yuyi, et al.
Veröffentlicht: (2025) -
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
von: Zhan, Runzhe, et al.
Veröffentlicht: (2025) -
Prefix Text as a Yarn: Eliciting Non-English Alignment in Foundation Language Model
von: Zhan, Runzhe, et al.
Veröffentlicht: (2024)