OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yitian, Cheng, Cheng, Sun, Yinan, Ling, Zi, Ge, Dongdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT
von: Ma, Chong, et al.
Veröffentlicht: (2023)
von: Ma, Chong, et al.
Veröffentlicht: (2023)
Deep Clustering via Gradual Community Detection
von: Cheng, Tianyu, et al.
Veröffentlicht: (2025)
von: Cheng, Tianyu, et al.
Veröffentlicht: (2025)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning Xiangqi Player with Monte Carlo Tree Search
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
von: Rosales, Rafael, et al.
Veröffentlicht: (2025)
von: Rosales, Rafael, et al.
Veröffentlicht: (2025)
Metaheuristics and Large Language Models Join Forces: Toward an Integrated Optimization Approach
von: Sartori, Camilo Chacón, et al.
Veröffentlicht: (2024)
von: Sartori, Camilo Chacón, et al.
Veröffentlicht: (2024)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning
von: Nayak, Nikhil Shivakumar, et al.
Veröffentlicht: (2025)
von: Nayak, Nikhil Shivakumar, et al.
Veröffentlicht: (2025)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
von: Wang, Youkang, et al.
Veröffentlicht: (2025)
von: Wang, Youkang, et al.
Veröffentlicht: (2025)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
The Meta-Prompting Protocol: Orchestrating LLMs via Adversarial Feedback Loops
von: Fu, Fanzhe
Veröffentlicht: (2025)
von: Fu, Fanzhe
Veröffentlicht: (2025)
In-Context Learning with Topological Information for Knowledge Graph Completion
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2024)
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2024)
TextClass Benchmark: A Continuous Elo Rating of LLMs in Social Sciences
von: González-Bustamante, Bastián
Veröffentlicht: (2024)
von: González-Bustamante, Bastián
Veröffentlicht: (2024)
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
von: Mohammadzadeh, Saeed, et al.
Veröffentlicht: (2025)
von: Mohammadzadeh, Saeed, et al.
Veröffentlicht: (2025)
Benchmarking LLMs in Political Content Text-Annotation: Proof-of-Concept with Toxicity and Incivility Data
von: González-Bustamante, Bastián
Veröffentlicht: (2024)
von: González-Bustamante, Bastián
Veröffentlicht: (2024)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
von: Das, Sourav
Veröffentlicht: (2026)
von: Das, Sourav
Veröffentlicht: (2026)
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
von: Yazdanjue, Navid, et al.
Veröffentlicht: (2025)
von: Yazdanjue, Navid, et al.
Veröffentlicht: (2025)
One Prompt is not Enough: Automated Construction of a Mixture-of-Expert Prompts
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
Towards Interpretable Soft Prompts
von: Patel, Oam, et al.
Veröffentlicht: (2025)
von: Patel, Oam, et al.
Veröffentlicht: (2025)
Observations on LLMs for Telecom Domain: Capabilities and Limitations
von: Soman, Sumit, et al.
Veröffentlicht: (2023)
von: Soman, Sumit, et al.
Veröffentlicht: (2023)
Achieving Distributive Justice in Federated Learning via Uncertainty Quantification
von: Carey, Alycia, et al.
Veröffentlicht: (2025)
von: Carey, Alycia, et al.
Veröffentlicht: (2025)
A Neural Affinity Framework for Abstract Reasoning: Diagnosing the Compositional Gap in Transformer Architectures via Procedural Task Taxonomy
von: Ingram, Miguel, et al.
Veröffentlicht: (2025)
von: Ingram, Miguel, et al.
Veröffentlicht: (2025)
DiTTO-LLM: Framework for Discovering Topic-based Technology Opportunities via Large Language Model
von: Kim, Wonyoung, et al.
Veröffentlicht: (2025)
von: Kim, Wonyoung, et al.
Veröffentlicht: (2025)
WebCanvas: Benchmarking Web Agents in Online Environments
von: Pan, Yichen, et al.
Veröffentlicht: (2024)
von: Pan, Yichen, et al.
Veröffentlicht: (2024)
When Emotional Stimuli meet Prompt Designing: An Auto-Prompt Graphical Paradigm
von: Ma, Chenggian, et al.
Veröffentlicht: (2024)
von: Ma, Chenggian, et al.
Veröffentlicht: (2024)
Compressible Softmax-Attended Language under Incompressible Attention
von: Lee, Wonsuk
Veröffentlicht: (2026)
von: Lee, Wonsuk
Veröffentlicht: (2026)
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
von: Gulati, Aryan, et al.
Veröffentlicht: (2025)
von: Gulati, Aryan, et al.
Veröffentlicht: (2025)
A Primer on Large Language Models and their Limitations
von: Johnson, Sandra, et al.
Veröffentlicht: (2024)
von: Johnson, Sandra, et al.
Veröffentlicht: (2024)
Large Language Models Report Subjective Experience Under Self-Referential Processing
von: Berg, Cameron, et al.
Veröffentlicht: (2025)
von: Berg, Cameron, et al.
Veröffentlicht: (2025)
Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
von: Liu, Xiaoou, et al.
Veröffentlicht: (2026)
von: Liu, Xiaoou, et al.
Veröffentlicht: (2026)
Improving Time Series Classification with Representation Soft Label Smoothing
von: Ma, Hengyi, et al.
Veröffentlicht: (2024)
von: Ma, Hengyi, et al.
Veröffentlicht: (2024)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
von: Abhishek, Alok, et al.
Veröffentlicht: (2026)
Retrieval Quality at Context Limit
von: McKinnon, Max
Veröffentlicht: (2025)
von: McKinnon, Max
Veröffentlicht: (2025)
Large-scale Urban Facility Location Selection with Knowledge-informed Reinforcement Learning
von: Su, Hongyuan, et al.
Veröffentlicht: (2024)
von: Su, Hongyuan, et al.
Veröffentlicht: (2024)
Regime Change Hypothesis: Foundations for Decoupled Dynamics in Neural Network Training
von: Pérez-Corral, Cristian, et al.
Veröffentlicht: (2026)
von: Pérez-Corral, Cristian, et al.
Veröffentlicht: (2026)
Enhancing Feature Selection and Interpretability in AI Regression Tasks Through Feature Attribution
von: Hinterleitner, Alexander, et al.
Veröffentlicht: (2024)
von: Hinterleitner, Alexander, et al.
Veröffentlicht: (2024)
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Dark LLMs: The Growing Threat of Unaligned AI Models
von: Fire, Michael, et al.
Veröffentlicht: (2025)
von: Fire, Michael, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT
von: Ma, Chong, et al.
Veröffentlicht: (2023) -
Deep Clustering via Gradual Community Detection
von: Cheng, Tianyu, et al.
Veröffentlicht: (2025) -
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
von: Wang, Zhen, et al.
Veröffentlicht: (2025) -
Deep Reinforcement Learning Xiangqi Player with Monte Carlo Tree Search
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025) -
Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
von: Rosales, Rafael, et al.
Veröffentlicht: (2025)