Saved in:
Bibliographic Details
Main Authors: Camuffo, Arnaldo, Gambardella, Alfonso, Kazemi, Saeid, Malachowski, Jakub, Pandey, Abhinav
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2601.02370
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917208999854080
author Camuffo, Arnaldo
Gambardella, Alfonso
Kazemi, Saeid
Malachowski, Jakub
Pandey, Abhinav
author_facet Camuffo, Arnaldo
Gambardella, Alfonso
Kazemi, Saeid
Malachowski, Jakub
Pandey, Abhinav
contents Large language models (LLMs) offer strategy researchers powerful tools for annotating text at scale, but treating LLM-generated labels as deterministic overlooks substantial instability. Grounded in content analysis and generalizability theory, we diagnose five variance sources: construct specification, interface effects, model preferences, output extraction, and system-level aggregation. Empirical demonstrations show that minor design choices-prompt phrasing, model selection-can shift outcomes by 12-85 percentage points. Such variance threatens not only reproducibility but econometric identification: annotation errors correlated with covariates bias parameter estimates regardless of average accuracy. We develop a variance-aware protocol specifying sampling budgets, aggregation rules, and reporting standards, and delineate scope conditions where LLM annotation should not be used. These contributions transform LLM-based annotation from ad hoc practice into auditable measurement infrastructure.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02370
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Variance-Aware LLM Annotation for Strategy Research: Sources, Diagnostics, and a Protocol for Reliable Measurement
Camuffo, Arnaldo
Gambardella, Alfonso
Kazemi, Saeid
Malachowski, Jakub
Pandey, Abhinav
Computers and Society
Computation and Language
68T07
C.4; I.2.6; I.2.7
Large language models (LLMs) offer strategy researchers powerful tools for annotating text at scale, but treating LLM-generated labels as deterministic overlooks substantial instability. Grounded in content analysis and generalizability theory, we diagnose five variance sources: construct specification, interface effects, model preferences, output extraction, and system-level aggregation. Empirical demonstrations show that minor design choices-prompt phrasing, model selection-can shift outcomes by 12-85 percentage points. Such variance threatens not only reproducibility but econometric identification: annotation errors correlated with covariates bias parameter estimates regardless of average accuracy. We develop a variance-aware protocol specifying sampling budgets, aggregation rules, and reporting standards, and delineate scope conditions where LLM annotation should not be used. These contributions transform LLM-based annotation from ad hoc practice into auditable measurement infrastructure.
title Variance-Aware LLM Annotation for Strategy Research: Sources, Diagnostics, and a Protocol for Reliable Measurement
topic Computers and Society
Computation and Language
68T07
C.4; I.2.6; I.2.7
url https://arxiv.org/abs/2601.02370