Evaluating Large Language Model Capabilities in Assessing Spatial Econometrics Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arbia, Giuseppe, Morandini, Luca, Nardelli, Vincenzo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913882730135552
author Arbia, Giuseppe
Morandini, Luca
Nardelli, Vincenzo
author_facet Arbia, Giuseppe
Morandini, Luca
Nardelli, Vincenzo
contents This paper investigates Large Language Models (LLMs) ability to assess the economic soundness and theoretical consistency of empirical findings in spatial econometrics. We created original and deliberately altered "counterfactual" summaries from 28 published papers (2005-2024), which were evaluated by a diverse set of LLMs. The LLMs provided qualitative assessments and structured binary classifications on variable choice, coefficient plausibility, and publication suitability. The results indicate that while LLMs can expertly assess the coherence of variable choices (with top models like GPT-4o achieving an overall F1 score of 0.87), their performance varies significantly when evaluating deeper aspects such as coefficient plausibility and overall publication suitability. The results further revealed that the choice of LLM, the specific characteristics of the paper and the interaction between these two factors significantly influence the accuracy of the assessment, particularly for nuanced judgments. These findings highlight LLMs' current strengths in assisting with initial, more surface-level checks and their limitations in performing comprehensive, deep economic reasoning, suggesting a potential assistive role in peer review that still necessitates robust human oversight.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06377
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Large Language Model Capabilities in Assessing Spatial Econometrics Research
Arbia, Giuseppe
Morandini, Luca
Nardelli, Vincenzo
Computers and Society
Machine Learning
Econometrics
Computation
This paper investigates Large Language Models (LLMs) ability to assess the economic soundness and theoretical consistency of empirical findings in spatial econometrics. We created original and deliberately altered "counterfactual" summaries from 28 published papers (2005-2024), which were evaluated by a diverse set of LLMs. The LLMs provided qualitative assessments and structured binary classifications on variable choice, coefficient plausibility, and publication suitability. The results indicate that while LLMs can expertly assess the coherence of variable choices (with top models like GPT-4o achieving an overall F1 score of 0.87), their performance varies significantly when evaluating deeper aspects such as coefficient plausibility and overall publication suitability. The results further revealed that the choice of LLM, the specific characteristics of the paper and the interaction between these two factors significantly influence the accuracy of the assessment, particularly for nuanced judgments. These findings highlight LLMs' current strengths in assisting with initial, more surface-level checks and their limitations in performing comprehensive, deep economic reasoning, suggesting a potential assistive role in peer review that still necessitates robust human oversight.
title Evaluating Large Language Model Capabilities in Assessing Spatial Econometrics Research
topic Computers and Society
Machine Learning
Econometrics
Computation
url https://arxiv.org/abs/2506.06377