Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Gupta, Himanshu, Verma, Shreyas, Anantheswaran, Ujjwala, Scaria, Kevin, Parmar, Mihir, Mishra, Swaroop, Baral, Chitta |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
di: Anantheswaran, Ujjwala, et al.
Pubblicazione: (2024)
di: Anantheswaran, Ujjwala, et al.
Pubblicazione: (2024)
TarGEN: Targeted Data Generation with Large Language Models
di: Gupta, Himanshu, et al.
Pubblicazione: (2023)
di: Gupta, Himanshu, et al.
Pubblicazione: (2023)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
di: Parmar, Mihir, et al.
Pubblicazione: (2022)
di: Parmar, Mihir, et al.
Pubblicazione: (2022)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
di: Patel, Nisarg, et al.
Pubblicazione: (2024)
di: Patel, Nisarg, et al.
Pubblicazione: (2024)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
di: Luo, Man, et al.
Pubblicazione: (2023)
di: Luo, Man, et al.
Pubblicazione: (2023)
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
di: Parmar, Mihir, et al.
Pubblicazione: (2025)
di: Parmar, Mihir, et al.
Pubblicazione: (2025)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
di: Tyagi, Nemika, et al.
Pubblicazione: (2024)
di: Tyagi, Nemika, et al.
Pubblicazione: (2024)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
di: Mukhopadhyay, Souradeep, et al.
Pubblicazione: (2025)
di: Mukhopadhyay, Souradeep, et al.
Pubblicazione: (2025)
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
di: Parmar, Mihir, et al.
Pubblicazione: (2025)
di: Parmar, Mihir, et al.
Pubblicazione: (2025)
ThinkTuning: Instilling Cognitive Reflections without Distillation
di: RRV, Aswin, et al.
Pubblicazione: (2025)
di: RRV, Aswin, et al.
Pubblicazione: (2025)
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
di: RRV, Aswin, et al.
Pubblicazione: (2026)
di: RRV, Aswin, et al.
Pubblicazione: (2026)
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
di: Mishra, Venkatesh, et al.
Pubblicazione: (2025)
di: Mishra, Venkatesh, et al.
Pubblicazione: (2025)
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
di: Handa, Divij, et al.
Pubblicazione: (2024)
di: Handa, Divij, et al.
Pubblicazione: (2024)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
di: Amjad, Husnain, et al.
Pubblicazione: (2026)
di: Amjad, Husnain, et al.
Pubblicazione: (2026)
Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
di: Verma, Pulkit, et al.
Pubblicazione: (2025)
di: Verma, Pulkit, et al.
Pubblicazione: (2025)
Towards Robust Mathematical Reasoning
di: Luong, Thang, et al.
Pubblicazione: (2025)
di: Luong, Thang, et al.
Pubblicazione: (2025)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
di: Mishra, Shubhra, et al.
Pubblicazione: (2024)
di: Mishra, Shubhra, et al.
Pubblicazione: (2024)
MMTABREAL: Real-World Benchmark for Multimodal Table Understanding
di: Titiya, Prasham, et al.
Pubblicazione: (2025)
di: Titiya, Prasham, et al.
Pubblicazione: (2025)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
di: He, Zheqi, et al.
Pubblicazione: (2024)
di: He, Zheqi, et al.
Pubblicazione: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
di: Yilmaz, Nilay, et al.
Pubblicazione: (2025)
di: Yilmaz, Nilay, et al.
Pubblicazione: (2025)
VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning
di: Jiang, Zhihuan, et al.
Pubblicazione: (2024)
di: Jiang, Zhihuan, et al.
Pubblicazione: (2024)
RoMath: A Mathematical Reasoning Benchmark in Romanian
di: Cosma, Adrian, et al.
Pubblicazione: (2024)
di: Cosma, Adrian, et al.
Pubblicazione: (2024)
Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness
di: Jayarao, Pratik, et al.
Pubblicazione: (2025)
di: Jayarao, Pratik, et al.
Pubblicazione: (2025)
Dissecting Physics Reasoning in Small Language Models: A Multi-Dimensional Analysis from an Educational Perspective
di: Scaria, Nicy, et al.
Pubblicazione: (2025)
di: Scaria, Nicy, et al.
Pubblicazione: (2025)
Automated Educational Question Generation at Different Bloom's Skill Levels using Large Language Models: Strategies and Evaluation
di: Scaria, Nicy, et al.
Pubblicazione: (2024)
di: Scaria, Nicy, et al.
Pubblicazione: (2024)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
di: Hong, Zijin, et al.
Pubblicazione: (2025)
di: Hong, Zijin, et al.
Pubblicazione: (2025)
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
di: Fabbri, Alexander R., et al.
Pubblicazione: (2025)
di: Fabbri, Alexander R., et al.
Pubblicazione: (2025)
Large Language Models Cannot Self-Correct Reasoning Yet
di: Huang, Jie, et al.
Pubblicazione: (2023)
di: Huang, Jie, et al.
Pubblicazione: (2023)
Sensitivity of Small Language Models to Fine-tuning Data Contamination
di: Scaria, Nicy, et al.
Pubblicazione: (2025)
di: Scaria, Nicy, et al.
Pubblicazione: (2025)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
di: Zhang, Yushun, et al.
Pubblicazione: (2026)
di: Zhang, Yushun, et al.
Pubblicazione: (2026)
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
di: Tran, Khanh-Tung, et al.
Pubblicazione: (2025)
di: Tran, Khanh-Tung, et al.
Pubblicazione: (2025)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
di: Zheng, Xiang, et al.
Pubblicazione: (2026)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
di: Siingh, Shikhhar, et al.
Pubblicazione: (2025)
di: Siingh, Shikhhar, et al.
Pubblicazione: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
di: Hao, Yuren, et al.
Pubblicazione: (2025)
di: Hao, Yuren, et al.
Pubblicazione: (2025)
Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning
di: Scaria, Nicy, et al.
Pubblicazione: (2026)
di: Scaria, Nicy, et al.
Pubblicazione: (2026)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
di: Saxon, Michael, et al.
Pubblicazione: (2024)
di: Saxon, Michael, et al.
Pubblicazione: (2024)
Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMs
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
di: Wang, Clinton J., et al.
Pubblicazione: (2025)
di: Wang, Clinton J., et al.
Pubblicazione: (2025)
EvalYaks: Instruction Tuning Datasets and LoRA Fine-tuned Models for Automated Scoring of CEFR B2 Speaking Assessment Transcripts
di: Scaria, Nicy, et al.
Pubblicazione: (2024)
di: Scaria, Nicy, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
di: Anantheswaran, Ujjwala, et al.
Pubblicazione: (2024) -
TarGEN: Targeted Data Generation with Large Language Models
di: Gupta, Himanshu, et al.
Pubblicazione: (2023) -
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
di: Parmar, Mihir, et al.
Pubblicazione: (2022) -
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
di: Patel, Nisarg, et al.
Pubblicazione: (2024) -
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
di: Luo, Man, et al.
Pubblicazione: (2023)