Scientific Text Analysis with Robots applied to observatory proposals

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jerabkova, T., Boffin, H. M. J., Patat, F., Dorigo, D., Sogni, F., Primas, F.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916310523314176
author Jerabkova, T.
Boffin, H. M. J.
Patat, F.
Dorigo, D.
Sogni, F.
Primas, F.
author_facet Jerabkova, T.
Boffin, H. M. J.
Patat, F.
Dorigo, D.
Sogni, F.
Primas, F.
contents To test the potential disruptive effect of Artificial Intelligence (AI) transformers (e.g., ChatGPT) and their associated Large Language Models on the time allocation process, both in proposal reviewing and grading, an experiment has been set-up at ESO for the P112 Call for Proposals. The experiment aims at raising awareness in the ESO community and build valuable knowledge by identifying what future steps ESO and other observatories might need to take to stay up to date with current technologies. We present here the results of the experiment, which may further be used to inform decision-makers regarding the use of AI in the proposal review process. We find that the ChatGPT-adjusted proposals tend to receive lower grades compared to the original proposals. Moreover, ChatGPT 3.5 can generally not be trusted in providing correct scientific references, while the most recent version makes a better, but far from perfect, job. We also studied how ChatGPT deals with assessing proposals. It does an apparent remarkable job at providing a summary of ESO proposals, although it doesn't do so good to identify weaknesses. When looking at how it evaluates proposals, however, it appears that ChatGPT systematically gives a higher mark than humans, and tends to prefer proposals written by itself.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02992
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scientific Text Analysis with Robots applied to observatory proposals
Jerabkova, T.
Boffin, H. M. J.
Patat, F.
Dorigo, D.
Sogni, F.
Primas, F.
Instrumentation and Methods for Astrophysics
Physics and Society
To test the potential disruptive effect of Artificial Intelligence (AI) transformers (e.g., ChatGPT) and their associated Large Language Models on the time allocation process, both in proposal reviewing and grading, an experiment has been set-up at ESO for the P112 Call for Proposals. The experiment aims at raising awareness in the ESO community and build valuable knowledge by identifying what future steps ESO and other observatories might need to take to stay up to date with current technologies. We present here the results of the experiment, which may further be used to inform decision-makers regarding the use of AI in the proposal review process. We find that the ChatGPT-adjusted proposals tend to receive lower grades compared to the original proposals. Moreover, ChatGPT 3.5 can generally not be trusted in providing correct scientific references, while the most recent version makes a better, but far from perfect, job. We also studied how ChatGPT deals with assessing proposals. It does an apparent remarkable job at providing a summary of ESO proposals, although it doesn't do so good to identify weaknesses. When looking at how it evaluates proposals, however, it appears that ChatGPT systematically gives a higher mark than humans, and tends to prefer proposals written by itself.
title Scientific Text Analysis with Robots applied to observatory proposals
topic Instrumentation and Methods for Astrophysics
Physics and Society
url https://arxiv.org/abs/2407.02992