ElecBench: a Power Dispatch Evaluation Benchmark for Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhou, Xiyuan, Zhao, Huan, Cheng, Yuheng, Cao, Yuji, Liang, Gaoqi, Liu, Guolong, Liu, Wenxuan, Xu, Yan, Zhao, Junhua |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
GAIA -- A Large Language Model for Advanced Power Dispatch
par: Cheng, Yuheng, et autres
Publié: (2024)
par: Cheng, Yuheng, et autres
Publié: (2024)
Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods
par: Cao, Yuji, et autres
Publié: (2024)
par: Cao, Yuji, et autres
Publié: (2024)
Coordinated Power Smoothing Control for Wind Storage Integrated System with Physics-informed Deep Reinforcement Learning
par: Wang, Shuyi, et autres
Publié: (2024)
par: Wang, Shuyi, et autres
Publié: (2024)
EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
par: Zhou, Xiyuan, et autres
Publié: (2025)
par: Zhou, Xiyuan, et autres
Publié: (2025)
Applying Large Language Models to Power Systems: Potential Security Threats
par: Ruan, Jiaqi, et autres
Publié: (2023)
par: Ruan, Jiaqi, et autres
Publié: (2023)
Power System Fault Diagnosis with Quantum Computing and Efficient Gate Decomposition
par: Fei, Xiang, et autres
Publié: (2024)
par: Fei, Xiang, et autres
Publié: (2024)
EngiAgent: Fully Connected Coordination of LLM Agents for Solving Open-ended Engineering Problems with Feasible Solutions
par: Zhou, Xiyuan, et autres
Publié: (2026)
par: Zhou, Xiyuan, et autres
Publié: (2026)
SI-Bench: Benchmarking Social Intelligence of Large Language Models in Human-to-Human Conversations
par: Huang, Shuai, et autres
Publié: (2025)
par: Huang, Shuai, et autres
Publié: (2025)
Carbon Disclosure Effect, Corporate Fundamentals, and Net-zero Emission Target: Evidence from China
par: Zhou, Xiyuan, et autres
Publié: (2025)
par: Zhou, Xiyuan, et autres
Publié: (2025)
Multi-Objective Large Language Model Unlearning
par: Pan, Zibin, et autres
Publié: (2024)
par: Pan, Zibin, et autres
Publié: (2024)
π–π Stacking‐Directed Crystallographic Orientation of Zn Elec‐trodeposition for Ultralong‐Life Anodes
par: Xuelong Liao, et autres
Publié: (2026)
par: Xuelong Liao, et autres
Publié: (2026)
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
par: Yang, Qian, et autres
Publié: (2024)
par: Yang, Qian, et autres
Publié: (2024)
TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
par: Xiong, Zixin, et autres
Publié: (2026)
par: Xiong, Zixin, et autres
Publié: (2026)
Ramen: Robust Test-Time Adaptation of Vision-Language Models with Active Sample Selection
par: Bao, Wenxuan, et autres
Publié: (2026)
par: Bao, Wenxuan, et autres
Publié: (2026)
MSC-Bench: Benchmarking and Analyzing Multi-Sensor Corruption for Driving Perception
par: Hao, Xiaoshuai, et autres
Publié: (2025)
par: Hao, Xiaoshuai, et autres
Publié: (2025)
Framework of Resilient Transmission Network Reconfiguration Considering Cyber-Attacks
par: Yang, Chao, et autres
Publié: (2024)
par: Yang, Chao, et autres
Publié: (2024)
QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models
par: Wu, Yao, et autres
Publié: (2026)
par: Wu, Yao, et autres
Publié: (2026)
Large Language Model Powered Automated Modeling and Optimization of Active Distribution Network Dispatch Problems
par: Yang, Xu, et autres
Publié: (2025)
par: Yang, Xu, et autres
Publié: (2025)
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
par: Wang, Hui, et autres
Publié: (2025)
par: Wang, Hui, et autres
Publié: (2025)
Disentangling Language and Culture for Evaluating Multilingual Large Language Models
par: Ying, Jiahao, et autres
Publié: (2025)
par: Ying, Jiahao, et autres
Publié: (2025)
PE-GPT: A Physics-Informed Interactive Large Language Model for Power Converter Modulation Design
par: Lin, Fanfan, et autres
Publié: (2024)
par: Lin, Fanfan, et autres
Publié: (2024)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
par: Cao, Jialun, et autres
Publié: (2024)
par: Cao, Jialun, et autres
Publié: (2024)
A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
par: Liu, Jie, et autres
Publié: (2024)
par: Liu, Jie, et autres
Publié: (2024)
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models
par: Guo, Zhicheng, et autres
Publié: (2024)
par: Guo, Zhicheng, et autres
Publié: (2024)
A Distributionally Robust Optimization Framework for Stochastic Assessment of Power System Flexibility in Economic Dispatch
par: Zhao, Xinyi, et autres
Publié: (2024)
par: Zhao, Xinyi, et autres
Publié: (2024)
Pardon? Evaluating Conversational Repair in Large Audio-Language Models
par: Huang, Shuanghong, et autres
Publié: (2026)
par: Huang, Shuanghong, et autres
Publié: (2026)
SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management
par: Guan, Shengyue, et autres
Publié: (2026)
par: Guan, Shengyue, et autres
Publié: (2026)
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
par: Yin, Qiyue, et autres
Publié: (2025)
par: Yin, Qiyue, et autres
Publié: (2025)
SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
par: Ying, Zonghao, et autres
Publié: (2024)
par: Ying, Zonghao, et autres
Publié: (2024)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
par: Guo, Xiaoyu, et autres
Publié: (2025)
par: Guo, Xiaoyu, et autres
Publié: (2025)
Reflection-Bench: Evaluating Epistemic Agency in Large Language Models
par: Li, Lingyu, et autres
Publié: (2024)
par: Li, Lingyu, et autres
Publié: (2024)
CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
par: Wang, Weida, et autres
Publié: (2025)
par: Wang, Weida, et autres
Publié: (2025)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
par: Liu, Xuxu, et autres
Publié: (2025)
par: Liu, Xuxu, et autres
Publié: (2025)
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
par: Hu, Jiliang, et autres
Publié: (2025)
par: Hu, Jiliang, et autres
Publié: (2025)
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
par: Xu, Bin, et autres
Publié: (2025)
par: Xu, Bin, et autres
Publié: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
par: Han, ZhaoYang, et autres
Publié: (2025)
par: Han, ZhaoYang, et autres
Publié: (2025)
VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language Model
par: Wang, Sibo, et autres
Publié: (2024)
par: Wang, Sibo, et autres
Publié: (2024)
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
par: Liu, Xin, et autres
Publié: (2023)
par: Liu, Xin, et autres
Publié: (2023)
EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models
par: Chen, Yuyan, et autres
Publié: (2024)
par: Chen, Yuyan, et autres
Publié: (2024)
AirQualityBench: A Realistic Evaluation Benchmark for Global Air Quality Forecasting
par: Xu, Xing, et autres
Publié: (2026)
par: Xu, Xing, et autres
Publié: (2026)
Documents similaires
-
GAIA -- A Large Language Model for Advanced Power Dispatch
par: Cheng, Yuheng, et autres
Publié: (2024) -
Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods
par: Cao, Yuji, et autres
Publié: (2024) -
Coordinated Power Smoothing Control for Wind Storage Integrated System with Physics-informed Deep Reinforcement Learning
par: Wang, Shuyi, et autres
Publié: (2024) -
EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
par: Zhou, Xiyuan, et autres
Publié: (2025) -
Applying Large Language Models to Power Systems: Potential Security Threats
par: Ruan, Jiaqi, et autres
Publié: (2023)