MathDuels: Evaluating LLMs as Problem Posers and Solvers
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Zhiqiu, Jin, Shibo, Arya, Shreya, Naik, Mayur |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Solver-Independent Automated Problem Formulation via LLMs for High-Cost Simulation-Driven Design
por: Li, Yuchen, et al.
Publicado: (2025)
por: Li, Yuchen, et al.
Publicado: (2025)
Program Structure Aware Precondition Generation
por: Dinella, Elizabeth, et al.
Publicado: (2023)
por: Dinella, Elizabeth, et al.
Publicado: (2023)
IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities
por: Li, Ziyang, et al.
Publicado: (2024)
por: Li, Ziyang, et al.
Publicado: (2024)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
por: Yang, Zheyuan, et al.
Publicado: (2025)
por: Yang, Zheyuan, et al.
Publicado: (2025)
Evaluation of Code LLMs on Geospatial Code Generation
por: Gramacki, Piotr, et al.
Publicado: (2024)
por: Gramacki, Piotr, et al.
Publicado: (2024)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
por: Wu, Jiarong, et al.
Publicado: (2025)
por: Wu, Jiarong, et al.
Publicado: (2025)
QLCoder: A Query Synthesizer For Static Analysis of Security Vulnerabilities
por: Wang, Claire, et al.
Publicado: (2025)
por: Wang, Claire, et al.
Publicado: (2025)
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
por: Du, Yongkang, et al.
Publicado: (2025)
por: Du, Yongkang, et al.
Publicado: (2025)
Showing LLM-Generated Code Selectively Based on Confidence of LLMs
por: Li, Jia, et al.
Publicado: (2024)
por: Li, Jia, et al.
Publicado: (2024)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
por: Naik, Atharva, et al.
Publicado: (2024)
por: Naik, Atharva, et al.
Publicado: (2024)
SimCT: A Simple Consistency Test Protocol in LLMs Development Lifecycle
por: Zhao, Fufangchen, et al.
Publicado: (2024)
por: Zhao, Fufangchen, et al.
Publicado: (2024)
Tool-Aware Planning in Contact Center AI: Evaluating LLMs through Lineage-Guided Query Decomposition
por: Nathan, Varun, et al.
Publicado: (2026)
por: Nathan, Varun, et al.
Publicado: (2026)
Rethinking Repetition Problems of LLMs in Code Generation
por: Dong, Yihong, et al.
Publicado: (2025)
por: Dong, Yihong, et al.
Publicado: (2025)
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?
por: Ahmad, Wasi Uddin, et al.
Publicado: (2025)
por: Ahmad, Wasi Uddin, et al.
Publicado: (2025)
Type-Aware Retrieval-Augmented Generation with Dependency Closure for Solver-Executable Industrial Optimization Modeling
por: Zhong, Y., et al.
Publicado: (2026)
por: Zhong, Y., et al.
Publicado: (2026)
AutoCode: LLMs as Problem Setters for Competitive Programming
por: Zhou, Shang, et al.
Publicado: (2025)
por: Zhou, Shang, et al.
Publicado: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
por: Chen, Zaoyu, et al.
Publicado: (2026)
por: Chen, Zaoyu, et al.
Publicado: (2026)
Multi-Programming Language Sandbox for LLMs
por: Dou, Shihan, et al.
Publicado: (2024)
por: Dou, Shihan, et al.
Publicado: (2024)
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
por: Li, Xiangyang, et al.
Publicado: (2025)
por: Li, Xiangyang, et al.
Publicado: (2025)
CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora
por: Li, Shangyu, et al.
Publicado: (2026)
por: Li, Shangyu, et al.
Publicado: (2026)
Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey
por: Wang, Junqiao, et al.
Publicado: (2024)
por: Wang, Junqiao, et al.
Publicado: (2024)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
por: Wang, Hanbin, et al.
Publicado: (2025)
por: Wang, Hanbin, et al.
Publicado: (2025)
LLMs in Mobile Apps: Practices, Challenges, and Opportunities
por: Hau, Kimberly, et al.
Publicado: (2025)
por: Hau, Kimberly, et al.
Publicado: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
por: Li, Jia, et al.
Publicado: (2024)
por: Li, Jia, et al.
Publicado: (2024)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
por: Du, Junjia, et al.
Publicado: (2025)
por: Du, Junjia, et al.
Publicado: (2025)
Understanding the Effectiveness of Large Language Models in Detecting Security Vulnerabilities
por: Khare, Avishree, et al.
Publicado: (2023)
por: Khare, Avishree, et al.
Publicado: (2023)
Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
por: Chen, Kexin, et al.
Publicado: (2024)
por: Chen, Kexin, et al.
Publicado: (2024)
LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
por: Xia, Yunhui, et al.
Publicado: (2025)
por: Xia, Yunhui, et al.
Publicado: (2025)
CodeScout: Contextual Problem Statement Enhancement for Software Agents
por: Suri, Manan, et al.
Publicado: (2026)
por: Suri, Manan, et al.
Publicado: (2026)
NLPerturbator: Studying the Robustness of Code LLMs to Natural Language Variations
por: Chen, Junkai, et al.
Publicado: (2024)
por: Chen, Junkai, et al.
Publicado: (2024)
Model Editing for LLMs4Code: How Far are We?
por: Li, Xiaopeng, et al.
Publicado: (2024)
por: Li, Xiaopeng, et al.
Publicado: (2024)
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
por: Moon, Seungjun, et al.
Publicado: (2023)
por: Moon, Seungjun, et al.
Publicado: (2023)
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
por: Mu, Wenhan, et al.
Publicado: (2025)
por: Mu, Wenhan, et al.
Publicado: (2025)
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
por: Iskander, Shadi, et al.
Publicado: (2024)
por: Iskander, Shadi, et al.
Publicado: (2024)
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
por: Salimian, Sina, et al.
Publicado: (2025)
por: Salimian, Sina, et al.
Publicado: (2025)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
por: Huang, Dong, et al.
Publicado: (2025)
por: Huang, Dong, et al.
Publicado: (2025)
RITFIS: Robust input testing framework for LLMs-based intelligent software
por: Xiao, Mingxuan, et al.
Publicado: (2024)
por: Xiao, Mingxuan, et al.
Publicado: (2024)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
por: Zhang, Chenchen, et al.
Publicado: (2025)
por: Zhang, Chenchen, et al.
Publicado: (2025)
Leveraging LLMs for Grammar Adaptation: A Study on Metamodel-Grammar Co-Evolution
por: Zhang, Weixing, et al.
Publicado: (2026)
por: Zhang, Weixing, et al.
Publicado: (2026)
Ejemplares similares
-
Solver-Independent Automated Problem Formulation via LLMs for High-Cost Simulation-Driven Design
por: Li, Yuchen, et al.
Publicado: (2025) -
Program Structure Aware Precondition Generation
por: Dinella, Elizabeth, et al.
Publicado: (2023) -
IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities
por: Li, Ziyang, et al.
Publicado: (2024) -
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
por: Yang, Zheyuan, et al.
Publicado: (2025) -
Evaluation of Code LLMs on Geospatial Code Generation
por: Gramacki, Piotr, et al.
Publicado: (2024)