A Comparison of Conversational Models and Humans in Answering Technical Questions: the Firefox Case

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Correia, Joao, Coutinho, Daniel, Castelluccio, Marco, Barbosa, Caio, de Mello, Rafael, Sarma, Anita, Garcia, Alessandro, Gerosa, Marco, Steinmacher, Igor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914114199093248
author Correia, Joao
Coutinho, Daniel
Castelluccio, Marco
Barbosa, Caio
de Mello, Rafael
Sarma, Anita
Garcia, Alessandro
Gerosa, Marco
Steinmacher, Igor
author_facet Correia, Joao
Coutinho, Daniel
Castelluccio, Marco
Barbosa, Caio
de Mello, Rafael
Sarma, Anita
Garcia, Alessandro
Gerosa, Marco
Steinmacher, Igor
contents The use of Large Language Models (LLMs) to support tasks in software development has steadily increased over recent years. From assisting developers in coding activities to providing conversational agents that answer newcomers' questions. In collaboration with the Mozilla Foundation, this study evaluates the effectiveness of Retrieval-Augmented Generation (RAG) in assisting developers within the Mozilla Firefox project. We conducted an empirical analysis comparing responses from human developers, a standard GPT model, and a GPT model enhanced with RAG, using real queries from Mozilla's developer chat rooms. To ensure a rigorous evaluation, Mozilla experts assessed the responses based on helpfulness, comprehensiveness, and conciseness. The results show that RAG-assisted responses were more comprehensive than human developers (62.50% to 54.17%) and almost as helpful (75.00% to 79.17%), suggesting RAG's potential to enhance developer assistance. However, the RAG responses were not as concise and often verbose. The results show the potential to apply RAG-based tools to Open Source Software (OSS) to minimize the load to core maintainers without losing answer quality. Toning down retrieval mechanisms and making responses even shorter in the future would enhance developer assistance in massive projects like Mozilla Firefox.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21933
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comparison of Conversational Models and Humans in Answering Technical Questions: the Firefox Case
Correia, Joao
Coutinho, Daniel
Castelluccio, Marco
Barbosa, Caio
de Mello, Rafael
Sarma, Anita
Garcia, Alessandro
Gerosa, Marco
Steinmacher, Igor
Software Engineering
Artificial Intelligence
The use of Large Language Models (LLMs) to support tasks in software development has steadily increased over recent years. From assisting developers in coding activities to providing conversational agents that answer newcomers' questions. In collaboration with the Mozilla Foundation, this study evaluates the effectiveness of Retrieval-Augmented Generation (RAG) in assisting developers within the Mozilla Firefox project. We conducted an empirical analysis comparing responses from human developers, a standard GPT model, and a GPT model enhanced with RAG, using real queries from Mozilla's developer chat rooms. To ensure a rigorous evaluation, Mozilla experts assessed the responses based on helpfulness, comprehensiveness, and conciseness. The results show that RAG-assisted responses were more comprehensive than human developers (62.50% to 54.17%) and almost as helpful (75.00% to 79.17%), suggesting RAG's potential to enhance developer assistance. However, the RAG responses were not as concise and often verbose. The results show the potential to apply RAG-based tools to Open Source Software (OSS) to minimize the load to core maintainers without losing answer quality. Toning down retrieval mechanisms and making responses even shorter in the future would enhance developer assistance in massive projects like Mozilla Firefox.
title A Comparison of Conversational Models and Humans in Answering Technical Questions: the Firefox Case
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2510.21933