| _version_ | 1866901781314797568 |
|---|---|
| author | Ishibashi, Ryuhei |
| author_facet | Ishibashi, Ryuhei |
| contents | <p>Recent large language models (LLMs) have been reported to achieve scores above the human average on standardized intelligence tests such as the WAIS-IV. These findings are often interpreted as evidence that AI systems possess human-like or even "superhuman" general intelligence. From a psychometric perspective, however, such interpretations ignore the valid range and design assumptions of the instruments being used.</p> <p>In this paper, we formalize three ways in which applying human IQ tests to high-performing AI systems constitutes a misuse of the measurement tool:<br>(1) Test information functions reveal that standard IQ tests provide negligible information in the high-ability range in which AI is purported to lie, making any between-model differences practically undetectable;<br>(2) Structural contamination of training data transforms nominally fluid-intelligence items into crystallized pattern-retrieval tasks, undermining construct validity through structural isomorphism; and<br>(3) The g-factor and positive-manifold assumptions underlying predictive interpretations of IQ scores fail for systems with jagged, non-human ability profiles.</p> <p>Using AI as a stress test for psychometric theory, we propose an explicit definition of valid measurement range based on item response theory and outline design principles for future "Psychometrics 2.0" capable of handling both biological and artificial agents.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_17718517 |
| institution | Zenodo |
| language | eng |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Out of Range: The Psychometric Mismatch Between Human IQ Tests and AI Evaluation Ishibashi, Ryuhei Large Language Models Psychometrics AI Evaluation IQ Tests Construct Validity Structural Contamination Item Response Theory Valid Range Artificial General Intelligence Psychometrics 2.0 <p>Recent large language models (LLMs) have been reported to achieve scores above the human average on standardized intelligence tests such as the WAIS-IV. These findings are often interpreted as evidence that AI systems possess human-like or even "superhuman" general intelligence. From a psychometric perspective, however, such interpretations ignore the valid range and design assumptions of the instruments being used.</p> <p>In this paper, we formalize three ways in which applying human IQ tests to high-performing AI systems constitutes a misuse of the measurement tool:<br>(1) Test information functions reveal that standard IQ tests provide negligible information in the high-ability range in which AI is purported to lie, making any between-model differences practically undetectable;<br>(2) Structural contamination of training data transforms nominally fluid-intelligence items into crystallized pattern-retrieval tasks, undermining construct validity through structural isomorphism; and<br>(3) The g-factor and positive-manifold assumptions underlying predictive interpretations of IQ scores fail for systems with jagged, non-human ability profiles.</p> <p>Using AI as a stress test for psychometric theory, we propose an explicit definition of valid measurement range based on item response theory and outline design principles for future "Psychometrics 2.0" capable of handling both biological and artificial agents.</p> |
| title | Out of Range: The Psychometric Mismatch Between Human IQ Tests and AI Evaluation |
| topic | Large Language Models Psychometrics AI Evaluation IQ Tests Construct Validity Structural Contamination Item Response Theory Valid Range Artificial General Intelligence Psychometrics 2.0 |
| url | https://doi.org/10.5281/zenodo.17718517 |