ResearchMath-14K: Scaling Research-Level Mathematics via Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Son, Guijin, Yi, Seungyeop, Gwak, Minju, Ko, Hyunwoo, Jang, Wongi, Yu, Youngjae |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
di: Son, Guijin, et al.
Pubblicazione: (2026)
di: Son, Guijin, et al.
Pubblicazione: (2026)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
di: Gwak, Minju, et al.
Pubblicazione: (2025)
di: Gwak, Minju, et al.
Pubblicazione: (2025)
Revisiting the UID Hypothesis in LLM Reasoning Traces
di: Gwak, Minju, et al.
Pubblicazione: (2025)
di: Gwak, Minju, et al.
Pubblicazione: (2025)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025)
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025)
Multi-Step Reasoning in Korean and the Emergent Mirage
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
Won: Establishing Best Practices for Korean Financial NLP
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
KAIO: A Collection of More Challenging Korean Questions
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
Controlling Language Confusion in Multilingual LLMs
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback
di: Son, Guijin, et al.
Pubblicazione: (2026)
di: Son, Guijin, et al.
Pubblicazione: (2026)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
di: Son, Guijin, et al.
Pubblicazione: (2026)
di: Son, Guijin, et al.
Pubblicazione: (2026)
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
di: Yan, Yibo, et al.
Pubblicazione: (2025)
di: Yan, Yibo, et al.
Pubblicazione: (2025)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
di: He, Zhiwei, et al.
Pubblicazione: (2025)
di: He, Zhiwei, et al.
Pubblicazione: (2025)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
di: Tang, Zhengyang, et al.
Pubblicazione: (2024)
di: Tang, Zhengyang, et al.
Pubblicazione: (2024)
Towards Lifelong Dialogue Agents via Timeline-based Memory Management
di: Ong, Kai Tzu-iunn, et al.
Pubblicazione: (2024)
di: Ong, Kai Tzu-iunn, et al.
Pubblicazione: (2024)
SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation
di: Wei, Hu, et al.
Pubblicazione: (2025)
di: Wei, Hu, et al.
Pubblicazione: (2025)
ESG Classification by Implicit Rule Learning via GPT-4
di: Yun, Hyo Jeong, et al.
Pubblicazione: (2024)
di: Yun, Hyo Jeong, et al.
Pubblicazione: (2024)
MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis
di: Yu, Zixiong, et al.
Pubblicazione: (2026)
di: Yu, Zixiong, et al.
Pubblicazione: (2026)
Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification
di: Jeong, Jinhong, et al.
Pubblicazione: (2026)
di: Jeong, Jinhong, et al.
Pubblicazione: (2026)
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
di: Luo, Haipeng, et al.
Pubblicazione: (2025)
di: Luo, Haipeng, et al.
Pubblicazione: (2025)
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
di: Liu, Wentao, et al.
Pubblicazione: (2024)
di: Liu, Wentao, et al.
Pubblicazione: (2024)
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
di: Chae, Hyungjoo, et al.
Pubblicazione: (2024)
di: Chae, Hyungjoo, et al.
Pubblicazione: (2024)
MathLearner: A Large Language Model Agent Framework for Learning to Solve Mathematical Problems
di: Xie, Wenbei, et al.
Pubblicazione: (2024)
di: Xie, Wenbei, et al.
Pubblicazione: (2024)
IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
di: Schmitt, Johannes, et al.
Pubblicazione: (2025)
di: Schmitt, Johannes, et al.
Pubblicazione: (2025)
PersonaMath: Boosting Mathematical Reasoning via Persona-Driven Data Augmentation
di: Luo, Jing, et al.
Pubblicazione: (2024)
di: Luo, Jing, et al.
Pubblicazione: (2024)
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?
di: Raimondi, Bianca, et al.
Pubblicazione: (2026)
di: Raimondi, Bianca, et al.
Pubblicazione: (2026)
MathEDU: Feedback Generation on Problem-Solving Processes for Mathematical Learning Support
di: Hsu, Wei-Ling, et al.
Pubblicazione: (2025)
di: Hsu, Wei-Ling, et al.
Pubblicazione: (2025)
Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
di: Lee, Seungbeen, et al.
Pubblicazione: (2025)
di: Lee, Seungbeen, et al.
Pubblicazione: (2025)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
di: Wang, Yiming, et al.
Pubblicazione: (2025)
di: Wang, Yiming, et al.
Pubblicazione: (2025)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
di: Shi, Wenhao, et al.
Pubblicazione: (2024)
di: Shi, Wenhao, et al.
Pubblicazione: (2024)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
di: Choi, Dasol, et al.
Pubblicazione: (2024)
di: Choi, Dasol, et al.
Pubblicazione: (2024)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
di: Dekoninck, Jasper, et al.
Pubblicazione: (2026)
di: Dekoninck, Jasper, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
di: Son, Guijin, et al.
Pubblicazione: (2026) -
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
di: Son, Guijin, et al.
Pubblicazione: (2025) -
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
di: Gwak, Minju, et al.
Pubblicazione: (2025) -
Revisiting the UID Hypothesis in LLM Reasoning Traces
di: Gwak, Minju, et al.
Pubblicazione: (2025) -
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025)