ElecBench: a Power Dispatch Evaluation Benchmark for Large Language Models

Xiyuan Zhou,Huan Zhao,Yuheng Cheng,Yuji Cao,Gaoqi Liang,Guolong Liu,Wenxuan Liu,Yan Xu,Junhua Zhao
2024-08-11
Abstract:In response to the urgent demand for grid stability and the complex challenges posed by renewable energy integration and electricity market dynamics, the power sector increasingly seeks innovative technological solutions. In this context, large language models (LLMs) have become a key technology to improve efficiency and promote intelligent progress in the power sector with their excellent natural language processing, logical reasoning, and generalization capabilities. Despite their potential, the absence of a performance evaluation benchmark for LLM in the power sector has limited the effective application of these technologies. Addressing this gap, our study introduces "ElecBench", an evaluation benchmark of LLMs within the power sector. ElecBench aims to overcome the shortcomings of existing evaluation benchmarks by providing comprehensive coverage of sector-specific scenarios, deepening the testing of professional knowledge, and enhancing decision-making precision. The framework categorizes scenarios into general knowledge and professional business, further divided into six core performance metrics: factuality, logicality, stability, security, fairness, and expressiveness, and is subdivided into 24 sub-metrics, offering profound insights into the capabilities and limitations of LLM applications in the power sector. To ensure transparency, we have made the complete test set public, evaluating the performance of eight LLMs across various scenarios and metrics. ElecBench aspires to serve as the standard benchmark for LLM applications in the power sector, supporting continuous updates of scenarios, metrics, and models to drive technological progress and application.
Artificial Intelligence
What problem does this paper attempt to address?
The problem this paper attempts to address is the lack of a benchmark framework for evaluating the performance of large language models (LLMs) in the field of power system operations. Although existing natural language processing (NLP) evaluation benchmarks such as GLUE, SuperGLUE, and SQuAD can assess individual capabilities, they do not adequately cover the specific needs and business scenarios of the power system operations field, particularly in handling professional technical and engineering tasks. Additionally, existing evaluation frameworks have shortcomings in dealing with simulated data in specific domains, which limits the performance evaluation of LLMs in power system operation tasks. To address these challenges, this study proposes an innovative evaluation framework called "ElecBench," aimed at comprehensively analyzing the general and specific performance of LLMs in power system operation tasks. By simulating detailed operational scenarios and their sub-scenarios, this framework can accurately assess the ability of LLMs to handle power system operation issues, ensuring comprehensive and in-depth evaluation. Specifically, the ElecBench framework includes six core performance indicators: factuality, logic, stability, fairness, safety, and expressiveness, further subdivided into 24 sub-indicators, providing profound insights into the application capabilities and limitations of LLMs in the power domain. The main contributions include: 1. Proposing a pioneering evaluation framework specifically for the power domain, covering six basic indicators and 24 secondary indicators, which is the first comprehensive framework for evaluating the performance of LLMs in the power domain. 2. Developing a method for generating test data and creating a dataset specifically for evaluating LLMs in power system operation challenges, which has been publicly released to promote further research in this field. 3. Conducting empirical tests to evaluate the performance of several cutting-edge LLMs, demonstrating the current application effects of LLMs in the power domain, and providing valuable insights and strategies for future development and practical applications.