OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xinyu, Zhang, Boxuan, Wan, Yuchen, Zhang, Lingling, Yao, YiXing, Wei, Bifan, Wu, Yaqiang, Liu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
$\textbf{AGT$^{AO}$}$: Robust and Stabilized LLM Unlearning via Adversarial Gating Training with Adaptive Orthogonality
by: Li, Pengyu, et al.
Published: (2026)
by: Li, Pengyu, et al.
Published: (2026)
QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
by: Fu, Weiping, et al.
Published: (2024)
by: Fu, Weiping, et al.
Published: (2024)
What Shapes a Creative Machine Mind? Comprehensively Benchmarking Creativity in Foundation Models
by: He, Zicong, et al.
Published: (2025)
by: He, Zicong, et al.
Published: (2025)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Solving Math Word Problems via Cooperative Reasoning induced Language Models
by: Zhu, Xinyu, et al.
Published: (2022)
by: Zhu, Xinyu, et al.
Published: (2022)
Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
by: Zheng, Shunfeng, et al.
Published: (2025)
by: Zheng, Shunfeng, et al.
Published: (2025)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
by: Fang, Meng, et al.
Published: (2024)
by: Fang, Meng, et al.
Published: (2024)
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and More
by: Huang, Muye, et al.
Published: (2026)
by: Huang, Muye, et al.
Published: (2026)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
LLM-OptiRA: LLM-Driven Optimization of Resource Allocation for Non-Convex Problems in Wireless Communications
by: Peng, Xinyue, et al.
Published: (2025)
by: Peng, Xinyue, et al.
Published: (2025)
Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator
by: Liu, Chengyuan, et al.
Published: (2024)
by: Liu, Chengyuan, et al.
Published: (2024)
SkillVerse : Assessing and Enhancing LLMs with Tree Evaluation
by: Tian, Yufei, et al.
Published: (2025)
by: Tian, Yufei, et al.
Published: (2025)
PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models
by: Wang, Yuwen, et al.
Published: (2026)
by: Wang, Yuwen, et al.
Published: (2026)
Collaborative Problem-Solving in an Optimization Game
by: Jeknic, Isidora, et al.
Published: (2025)
by: Jeknic, Isidora, et al.
Published: (2025)
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
HAD: HAllucination Detection Language Models Based on a Comprehensive Hallucination Taxonomy
by: Xu, Fan, et al.
Published: (2025)
by: Xu, Fan, et al.
Published: (2025)
Physics Reasoner: Knowledge-Augmented Reasoning for Solving Physics Problems with Large Language Models
by: Pang, Xinyu, et al.
Published: (2024)
by: Pang, Xinyu, et al.
Published: (2024)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
by: Zhang, Renrui, et al.
Published: (2024)
by: Zhang, Renrui, et al.
Published: (2024)
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
by: An, Wenbin, et al.
Published: (2025)
by: An, Wenbin, et al.
Published: (2025)
Shakespearean Sparks: The Dance of Hallucination and Creativity in LLMs' Decoding Layers
by: He, Zicong, et al.
Published: (2025)
by: He, Zicong, et al.
Published: (2025)
Solving General Natural-Language-Description Optimization Problems with Large Language Models
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
AERO: Autonomous Evolutionary Reasoning Optimization via Endogenous Dual-Loop Feedback
by: Gao, Zhitao, et al.
Published: (2026)
by: Gao, Zhitao, et al.
Published: (2026)
Is Table Retrieval a Solved Problem? Exploring Join-Aware Multi-Table Retrieval
by: Chen, Peter Baile, et al.
Published: (2024)
by: Chen, Peter Baile, et al.
Published: (2024)
OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical Modeling
by: Bouscary, Maxime, et al.
Published: (2025)
by: Bouscary, Maxime, et al.
Published: (2025)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
by: Gomaa, Amr, et al.
Published: (2025)
by: Gomaa, Amr, et al.
Published: (2025)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Self-consistent Reasoning For Solving Math Word Problems
by: Xiong, Jing, et al.
Published: (2022)
by: Xiong, Jing, et al.
Published: (2022)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
by: Wu, Jiarong, et al.
Published: (2025)
by: Wu, Jiarong, et al.
Published: (2025)
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal Assistants
by: Qin, Lixiong, et al.
Published: (2025)
by: Qin, Lixiong, et al.
Published: (2025)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
by: Tong, Yuxuan, et al.
Published: (2024)
by: Tong, Yuxuan, et al.
Published: (2024)
Learning from Semi-Factuals: A Debiased and Semantic-Aware Framework for Generalized Relation Discovery
by: Wang, Jiaxin, et al.
Published: (2024)
by: Wang, Jiaxin, et al.
Published: (2024)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
by: Zhang, Yushun, et al.
Published: (2026)
by: Zhang, Yushun, et al.
Published: (2026)
EmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding
by: Li, Ao, et al.
Published: (2024)
by: Li, Ao, et al.
Published: (2024)
AwareCompiler: Agentic Context-Aware Compiler Optimization via a Synergistic Knowledge-Data Driven Framework
by: Lin, Hongyu, et al.
Published: (2025)
by: Lin, Hongyu, et al.
Published: (2025)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
by: Chen, Ziyang, et al.
Published: (2026)
by: Chen, Ziyang, et al.
Published: (2026)
Linear Reasoning vs. Proof by Cases: Obstacles for Large Language Models in FOL Problem Solving
by: Ji, Yuliang, et al.
Published: (2026)
by: Ji, Yuliang, et al.
Published: (2026)
Similar Items
-
Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving
by: Zhang, Xinyu, et al.
Published: (2026) -
$\textbf{AGT$^{AO}$}$: Robust and Stabilized LLM Unlearning via Adversarial Gating Training with Adaptive Orthogonality
by: Li, Pengyu, et al.
Published: (2026) -
QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
by: Fu, Weiping, et al.
Published: (2024) -
What Shapes a Creative Machine Mind? Comprehensively Benchmarking Creativity in Foundation Models
by: He, Zicong, et al.
Published: (2025) -
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)