ProOPF: Benchmarking and Improving LLMs for Professional-Grade Power Systems Optimization Modeling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Shen, Chao, Guo, Zihan, Wan, Xu, Yang, Zhenghao, Zhang, Yifan, Huang, Wengi, Song, Jie, Zhang, Zongyan, Sun, Mingyang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
par: Ai, Jiaxin, et autres
Publié: (2025)
par: Ai, Jiaxin, et autres
Publié: (2025)
StepGrade: Grading Programming Assignments with Context-Aware LLMs
par: Akyash, Mohammad, et autres
Publié: (2025)
par: Akyash, Mohammad, et autres
Publié: (2025)
DPO-F+: Aligning Code Repair Feedback with Developers' Preferences
par: Fang, Zihan, et autres
Publié: (2025)
par: Fang, Zihan, et autres
Publié: (2025)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
par: Xu, Xiangzhe, et autres
Publié: (2024)
par: Xu, Xiangzhe, et autres
Publié: (2024)
Rethinking Kernel Program Repair: Benchmarking and Enhancing LLMs with RGym
par: Shehada, Kareem, et autres
Publié: (2025)
par: Shehada, Kareem, et autres
Publié: (2025)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
par: Aleithan, Reem, et autres
Publié: (2024)
par: Aleithan, Reem, et autres
Publié: (2024)
UCRBench: Benchmarking LLMs on Use Case Recovery
par: Xiao, Shuyuan, et autres
Publié: (2025)
par: Xiao, Shuyuan, et autres
Publié: (2025)
Ising-based Test Optimization and Benchmarking
par: Yang, Yige, et autres
Publié: (2026)
par: Yang, Yige, et autres
Publié: (2026)
LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code
par: Guo, Liwei, et autres
Publié: (2025)
par: Guo, Liwei, et autres
Publié: (2025)
DePro: Understanding the Role of LLMs in Debugging Competitive Programming Code
par: Parvez, Nabiha, et autres
Publié: (2026)
par: Parvez, Nabiha, et autres
Publié: (2026)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
par: Wu, Jie JW, et autres
Publié: (2024)
par: Wu, Jie JW, et autres
Publié: (2024)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
par: Guo, Hanyang, et autres
Publié: (2025)
par: Guo, Hanyang, et autres
Publié: (2025)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
par: Huang, Dong, et autres
Publié: (2025)
par: Huang, Dong, et autres
Publié: (2025)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
par: Ouyang, Shuyin, et autres
Publié: (2025)
par: Ouyang, Shuyin, et autres
Publié: (2025)
LibRec: Benchmarking Retrieval-Augmented LLMs for Library Migration Recommendations
par: Han, Junxiao, et autres
Publié: (2025)
par: Han, Junxiao, et autres
Publié: (2025)
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
par: Feng, Dylan, et autres
Publié: (2026)
par: Feng, Dylan, et autres
Publié: (2026)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
par: Hu, Ruida, et autres
Publié: (2025)
par: Hu, Ruida, et autres
Publié: (2025)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
par: Chen, Yeheng, et autres
Publié: (2026)
par: Chen, Yeheng, et autres
Publié: (2026)
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
par: Ahmed, Md Basim Uddin, et autres
Publié: (2025)
par: Ahmed, Md Basim Uddin, et autres
Publié: (2025)
Production-Grade AI Coding System for Client-Side Development
par: Wang, Ruihan, et autres
Publié: (2026)
par: Wang, Ruihan, et autres
Publié: (2026)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
par: Jing, Huihao, et autres
Publié: (2026)
par: Jing, Huihao, et autres
Publié: (2026)
Benchmarking and Revisiting Code Generation Assessment: A Mutation-Based Approach
par: Wang, Longtian, et autres
Publié: (2025)
par: Wang, Longtian, et autres
Publié: (2025)
It's LIT! Reliability-Optimized LLMs with Inspectable Tools
par: Zhang, Ruixin, et autres
Publié: (2025)
par: Zhang, Ruixin, et autres
Publié: (2025)
Model-Assisted and Human-Guided: Perceptions and Practices of Software Professionals Using LLMs for Coding
par: Santos, Italo, et autres
Publié: (2025)
par: Santos, Italo, et autres
Publié: (2025)
Efficient DNN-Powered Software with Fair Sparse Models
par: Gao, Xuanqi, et autres
Publié: (2024)
par: Gao, Xuanqi, et autres
Publié: (2024)
Log Parsing using LLMs with Self-Generated In-Context Learning and Self-Correction
par: Wu, Yifan, et autres
Publié: (2024)
par: Wu, Yifan, et autres
Publié: (2024)
Boosting LLMs for Mutation Generation
par: Wang, Bo, et autres
Publié: (2026)
par: Wang, Bo, et autres
Publié: (2026)
Improving Code Translation with Syntax-Guided and Semantic-aware Preference Optimization
par: Wu, Yuhan, et autres
Publié: (2026)
par: Wu, Yuhan, et autres
Publié: (2026)
CLAP: Learning Transferable Binary Code Representations with Natural Language Supervision
par: Wang, Hao, et autres
Publié: (2024)
par: Wang, Hao, et autres
Publié: (2024)
Smart but Costly? Benchmarking LLMs on Functional Accuracy and Energy Efficiency
par: Mehditabar, Mohammadjavad, et autres
Publié: (2025)
par: Mehditabar, Mohammadjavad, et autres
Publié: (2025)
Unlocking the Power of Numbers: Log Compression via Numeric Token Parsing
par: Yu, Siyu, et autres
Publié: (2024)
par: Yu, Siyu, et autres
Publié: (2024)
ZeroFalse: Improving Precision in Static Analysis with LLMs
par: Iranmanesh, Mohsen, et autres
Publié: (2025)
par: Iranmanesh, Mohsen, et autres
Publié: (2025)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
par: Yu, Zhaojian, et autres
Publié: (2024)
par: Yu, Zhaojian, et autres
Publié: (2024)
COFFE: A Code Efficiency Benchmark for Code Generation
par: Peng, Yun, et autres
Publié: (2025)
par: Peng, Yun, et autres
Publié: (2025)
SmartLLMs Scheduler: A Framework for Cost-Effective LLMs Utilization
par: Liu, Yueyue, et autres
Publié: (2025)
par: Liu, Yueyue, et autres
Publié: (2025)
Rethinking the Evaluation of Microservice RCA with a Fault Propagation-Aware Benchmark
par: Fang, Aoyang, et autres
Publié: (2025)
par: Fang, Aoyang, et autres
Publié: (2025)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
par: Du, Junjia, et autres
Publié: (2025)
par: Du, Junjia, et autres
Publié: (2025)
Extracting Formal Specifications from Documents Using LLMs for Automated Testing
par: Li, Hui, et autres
Publié: (2025)
par: Li, Hui, et autres
Publié: (2025)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
par: Liu, Steven, et autres
Publié: (2026)
par: Liu, Steven, et autres
Publié: (2026)
EditFlow: Benchmarking and Optimizing Code Edit Recommendation Systems via Reconstruction of Developer Flows
par: Liu, Chenyan, et autres
Publié: (2026)
par: Liu, Chenyan, et autres
Publié: (2026)
Documents similaires
-
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
par: Ai, Jiaxin, et autres
Publié: (2025) -
StepGrade: Grading Programming Assignments with Context-Aware LLMs
par: Akyash, Mohammad, et autres
Publié: (2025) -
DPO-F+: Aligning Code Repair Feedback with Developers' Preferences
par: Fang, Zihan, et autres
Publié: (2025) -
ProSec: Fortifying Code LLMs with Proactive Security Alignment
par: Xu, Xiangzhe, et autres
Publié: (2024) -
Rethinking Kernel Program Repair: Benchmarking and Enhancing LLMs with RGym
par: Shehada, Kareem, et autres
Publié: (2025)