Max It or Miss It: Benchmarking LLM On Solving Extremal Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Binxin, Han, Jingjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
by: Singhi, Nishad, et al.
Published: (2025)
by: Singhi, Nishad, et al.
Published: (2025)
Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting
by: Lingo, Ryan, et al.
Published: (2024)
by: Lingo, Ryan, et al.
Published: (2024)
QuantumBench: A Benchmark for Quantum Problem Solving
by: Minami, Shunya, et al.
Published: (2025)
by: Minami, Shunya, et al.
Published: (2025)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
by: Do, Heejin, et al.
Published: (2026)
by: Do, Heejin, et al.
Published: (2026)
AcademicEval: Live Long-Context LLM Benchmark
by: Zhang, Haozhen, et al.
Published: (2025)
by: Zhang, Haozhen, et al.
Published: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
Graph of Thoughts: Solving Elaborate Problems with Large Language Models
by: Besta, Maciej, et al.
Published: (2023)
by: Besta, Maciej, et al.
Published: (2023)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
by: Sun, Yifan, et al.
Published: (2025)
by: Sun, Yifan, et al.
Published: (2025)
Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns?
by: Kao, Kuei-Chun, et al.
Published: (2024)
by: Kao, Kuei-Chun, et al.
Published: (2024)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
by: Walder, Christian, et al.
Published: (2025)
by: Walder, Christian, et al.
Published: (2025)
Boosting of Thoughts: Trial-and-Error Problem Solving with Large Language Models
by: Chen, Sijia, et al.
Published: (2024)
by: Chen, Sijia, et al.
Published: (2024)
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
by: MiniMax, et al.
Published: (2026)
by: MiniMax, et al.
Published: (2026)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2025)
by: Li, Dawei, et al.
Published: (2025)
Rethinking the Potential of Multimodality in Collaborative Problem Solving Diagnosis with Large Language Models
by: Wong, K., et al.
Published: (2025)
by: Wong, K., et al.
Published: (2025)
Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
by: Mahran, Mariam, et al.
Published: (2025)
by: Mahran, Mariam, et al.
Published: (2025)
Deconfounded Causality-aware Parameter-Efficient Fine-Tuning for Problem-Solving Improvement of LLMs
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
Can Generative AI Solve Your In-Context Learning Problem? A Martingale Perspective
by: Jesson, Andrew, et al.
Published: (2024)
by: Jesson, Andrew, et al.
Published: (2024)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
by: Wang, Xiaoxuan, et al.
Published: (2023)
by: Wang, Xiaoxuan, et al.
Published: (2023)
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
by: Potamitis, Nearchos, et al.
Published: (2025)
by: Potamitis, Nearchos, et al.
Published: (2025)
Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
by: Wang, Zhiwei, et al.
Published: (2024)
by: Wang, Zhiwei, et al.
Published: (2024)
Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?
by: Mittal, Chinmay, et al.
Published: (2024)
by: Mittal, Chinmay, et al.
Published: (2024)
DataSciBench: An LLM Agent Benchmark for Data Science
by: Zhang, Dan, et al.
Published: (2025)
by: Zhang, Dan, et al.
Published: (2025)
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
by: Zhao, Bingchen, et al.
Published: (2025)
by: Zhao, Bingchen, et al.
Published: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
by: White, Colin, et al.
Published: (2024)
by: White, Colin, et al.
Published: (2024)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025)
by: Kuzmanko, Jonathan
Published: (2025)
MarginSel : Max-Margin Demonstration Selection for LLMs
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
Incorporating Graph Attention Mechanism into Geometric Problem Solving Based on Deep Reinforcement Learning
by: Zhong, Xiuqin, et al.
Published: (2024)
by: Zhong, Xiuqin, et al.
Published: (2024)
Fleet of Agents: Coordinated Problem Solving with Large Language Models
by: Klein, Lars, et al.
Published: (2024)
by: Klein, Lars, et al.
Published: (2024)
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
by: Chen, Hongzheng, et al.
Published: (2025)
by: Chen, Hongzheng, et al.
Published: (2025)
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
by: Ni, Jinjie, et al.
Published: (2024)
by: Ni, Jinjie, et al.
Published: (2024)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
Solving the Content Gap in Roblox Game Recommendations: LLM-Based Profile Generation and Reranking
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models
by: Era, Jalisha Jashim, et al.
Published: (2025)
by: Era, Jalisha Jashim, et al.
Published: (2025)
Similar Items
-
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
by: Singhi, Nishad, et al.
Published: (2025) -
Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting
by: Lingo, Ryan, et al.
Published: (2024) -
QuantumBench: A Benchmark for Quantum Problem Solving
by: Minami, Shunya, et al.
Published: (2025) -
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
by: Do, Heejin, et al.
Published: (2026) -
AcademicEval: Live Long-Context LLM Benchmark
by: Zhang, Haozhen, et al.
Published: (2025)