Studying and Benchmarking Large Language Models For Log Level Suggestion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heng, Yi Wen, Ma, Zeyang, Li, Zhenhao, Kim, Dong Jae, Tse-Hsun, Chen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914970009075712
author Heng, Yi Wen
Ma, Zeyang
Li, Zhenhao
Kim, Dong Jae
Tse-Hsun
Chen
author_facet Heng, Yi Wen
Ma, Zeyang
Li, Zhenhao
Kim, Dong Jae
Tse-Hsun
Chen
contents Large Language Models (LLMs) have become a focal point of research across various domains, including software engineering, where their capabilities are increasingly leveraged. Recent studies have explored the integration of LLMs into software development tools and frameworks, revealing their potential to enhance performance in text and code-related tasks. Log level is a key part of a logging statement that allows software developers control the information recorded during system runtime. Given that log messages often mix natural language with code-like variables, LLMs' language translation abilities could be applied to determine the suitable verbosity level for logging statements. In this paper, we undertake a detailed empirical analysis to investigate the impact of characteristics and learning paradigms on the performance of 12 open-source LLMs in log level suggestion. We opted for open-source models because they enable us to utilize in-house code while effectively protecting sensitive information and maintaining data security. We examine several prompting strategies, including Zero-shot, Few-shot, and fine-tuning techniques, across different LLMs to identify the most effective combinations for accurate log level suggestions. Our research is supported by experiments conducted on 9 large-scale Java systems. The results indicate that although smaller LLMs can perform effectively with appropriate instruction and suitable techniques, there is still considerable potential for improvement in their ability to suggest log levels.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08499
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Studying and Benchmarking Large Language Models For Log Level Suggestion
Heng, Yi Wen
Ma, Zeyang
Li, Zhenhao
Kim, Dong Jae
Tse-Hsun
Chen
Software Engineering
Large Language Models (LLMs) have become a focal point of research across various domains, including software engineering, where their capabilities are increasingly leveraged. Recent studies have explored the integration of LLMs into software development tools and frameworks, revealing their potential to enhance performance in text and code-related tasks. Log level is a key part of a logging statement that allows software developers control the information recorded during system runtime. Given that log messages often mix natural language with code-like variables, LLMs' language translation abilities could be applied to determine the suitable verbosity level for logging statements. In this paper, we undertake a detailed empirical analysis to investigate the impact of characteristics and learning paradigms on the performance of 12 open-source LLMs in log level suggestion. We opted for open-source models because they enable us to utilize in-house code while effectively protecting sensitive information and maintaining data security. We examine several prompting strategies, including Zero-shot, Few-shot, and fine-tuning techniques, across different LLMs to identify the most effective combinations for accurate log level suggestions. Our research is supported by experiments conducted on 9 large-scale Java systems. The results indicate that although smaller LLMs can perform effectively with appropriate instruction and suitable techniques, there is still considerable potential for improvement in their ability to suggest log levels.
title Studying and Benchmarking Large Language Models For Log Level Suggestion
topic Software Engineering
url https://arxiv.org/abs/2410.08499