Trust, But Verify: An Empirical Evaluation of AI-Generated Code for SDN Controllers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Soares, Felipe Avencourt, Franco, Muriel F., Scheid, Eder J., Granville, Lisandro Z.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914109479452672
author Soares, Felipe Avencourt
Franco, Muriel F.
Scheid, Eder J.
Granville, Lisandro Z.
author_facet Soares, Felipe Avencourt
Franco, Muriel F.
Scheid, Eder J.
Granville, Lisandro Z.
contents Generative Artificial Intelligence (AI) tools have been used to generate human-like content across multiple domains (e.g., sound, image, text, and programming). However, their reliability in terms of correctness and functionality in novel contexts such as programmable networks remains unclear. Hence, this paper presents an empirical evaluation of the source code of a POX controller generated by different AI tools, namely ChatGPT, Copilot, DeepSeek, and BlackBox.ai. To evaluate such a code, three networking tasks of increasing complexity were defined and for each task, zero-shot and few-shot prompting techniques were input to the tools. Next, the output code was tested in emulated network topologies with Mininet and analyzed according to functionality, correctness, and the need for manual fixes. Results show that all evaluated models can produce functional controllers. However, ChatGPT and DeepSeek exhibited higher consistency and code quality, while Copilot and BlackBox.ai required more adjustments.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20703
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trust, But Verify: An Empirical Evaluation of AI-Generated Code for SDN Controllers
Soares, Felipe Avencourt
Franco, Muriel F.
Scheid, Eder J.
Granville, Lisandro Z.
Networking and Internet Architecture
Generative Artificial Intelligence (AI) tools have been used to generate human-like content across multiple domains (e.g., sound, image, text, and programming). However, their reliability in terms of correctness and functionality in novel contexts such as programmable networks remains unclear. Hence, this paper presents an empirical evaluation of the source code of a POX controller generated by different AI tools, namely ChatGPT, Copilot, DeepSeek, and BlackBox.ai. To evaluate such a code, three networking tasks of increasing complexity were defined and for each task, zero-shot and few-shot prompting techniques were input to the tools. Next, the output code was tested in emulated network topologies with Mininet and analyzed according to functionality, correctness, and the need for manual fixes. Results show that all evaluated models can produce functional controllers. However, ChatGPT and DeepSeek exhibited higher consistency and code quality, while Copilot and BlackBox.ai required more adjustments.
title Trust, But Verify: An Empirical Evaluation of AI-Generated Code for SDN Controllers
topic Networking and Internet Architecture
url https://arxiv.org/abs/2510.20703