Adapting Large Language Models to Log Analysis with Interpretable Domain Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Yuhe, Liu, Yilun, Yao, Feiyu, He, Minggui, Tao, Shimin, Zhao, Xiaofeng, Chang, Su, Yang, Xinhua, Meng, Weibin, Xie, Yuming, Chen, Boxing, Zhang, Shenglin, Sun, Yongqian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912553929539584
author Ji, Yuhe
Liu, Yilun
Yao, Feiyu
He, Minggui
Tao, Shimin
Zhao, Xiaofeng
Chang, Su
Yang, Xinhua
Meng, Weibin
Xie, Yuming
Chen, Boxing
Zhang, Shenglin
Sun, Yongqian
author_facet Ji, Yuhe
Liu, Yilun
Yao, Feiyu
He, Minggui
Tao, Shimin
Zhao, Xiaofeng
Chang, Su
Yang, Xinhua
Meng, Weibin
Xie, Yuming
Chen, Boxing
Zhang, Shenglin
Sun, Yongqian
contents Log analysis represents a critical sub-domain within AI applications that facilitates automatic approaches to fault and error management of large-scaled software systems, saving labors of traditional manual methods. While existing solutions using large language models (LLMs) show promise, they are limited by a significant domain gap between natural and log languages (the latter contains rich domain-specific tokens such as status codes, IP addresses, resource pathes), which restricts their effectiveness in real-world applications. However, directly adapting general-purpose LLMs to log analysis using raw logs may degrade their performance due to inconsistent token distribution. In this paper, we present a domain adaptation approach that addresses these limitations by integrating interpretable domain knowledge into open-source LLMs through continual pre-training (CPT), which bridges this domain gap by adapting LLMs on interpretable natural texts with log knowledge (instead of raw logs) to reduce distribution discrepancy. To achieve this, we developed NLPLog, a comprehensive dataset containing over 250,000 question-answer pairs on log-related knowledge. Our resulting model, SuperLog, achieves the best performance across four log analysis tasks, with an average accuracy improvement of 12.01% over the second-best model. Ablation study also suggests advantages of domain adaption using interpretable log knowledge over using raw logs.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01377
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adapting Large Language Models to Log Analysis with Interpretable Domain Knowledge
Ji, Yuhe
Liu, Yilun
Yao, Feiyu
He, Minggui
Tao, Shimin
Zhao, Xiaofeng
Chang, Su
Yang, Xinhua
Meng, Weibin
Xie, Yuming
Chen, Boxing
Zhang, Shenglin
Sun, Yongqian
Computation and Language
Software Engineering
Log analysis represents a critical sub-domain within AI applications that facilitates automatic approaches to fault and error management of large-scaled software systems, saving labors of traditional manual methods. While existing solutions using large language models (LLMs) show promise, they are limited by a significant domain gap between natural and log languages (the latter contains rich domain-specific tokens such as status codes, IP addresses, resource pathes), which restricts their effectiveness in real-world applications. However, directly adapting general-purpose LLMs to log analysis using raw logs may degrade their performance due to inconsistent token distribution. In this paper, we present a domain adaptation approach that addresses these limitations by integrating interpretable domain knowledge into open-source LLMs through continual pre-training (CPT), which bridges this domain gap by adapting LLMs on interpretable natural texts with log knowledge (instead of raw logs) to reduce distribution discrepancy. To achieve this, we developed NLPLog, a comprehensive dataset containing over 250,000 question-answer pairs on log-related knowledge. Our resulting model, SuperLog, achieves the best performance across four log analysis tasks, with an average accuracy improvement of 12.01% over the second-best model. Ablation study also suggests advantages of domain adaption using interpretable log knowledge over using raw logs.
title Adapting Large Language Models to Log Analysis with Interpretable Domain Knowledge
topic Computation and Language
Software Engineering
url https://arxiv.org/abs/2412.01377