30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Bernal Díaz, Víctor Cristóbal
Format: Recurso digital
Sprache:Spanisch
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901750034726912
author Bernal Díaz, Víctor Cristóbal
author_facet Bernal Díaz, Víctor Cristóbal
contents <h3><strong>RESUMEN EJECUTIVO:</strong></h3> <p>Este documento presenta la aplicación completa del protocolo RFC-EVAL-001 v1.1 para evaluar capacidades cognitivas complejas en seis sistemas de IA comerciales contemporáneos. Incluye 36 evaluaciones cruzadas ciegas, perfiles Ω dimensionales (Cm, Ht, Mm, Fa), clasificaciones NOVA, herramientas RFC-RC/RFC-CONEX para diseño de colaboraciones, y el dataset completo con código de replicación.</p> <p>Sobre la V2.1: Este documento ha sido sometido a un proceso de verificación metodológica independiente. Durante dicho proceso se identificaron y corrigieron errores en el cálculo de las puntuaciones Gamma (γ) y en el reporte de fiabilidad inter-evaluador presentes en una versión preliminar anterior.</p> <p>Sobre esta versión 2.2: Estado: Protocolo validado, datos verificados.</p> <p><em><strong>Benchmarking Complex Cognitive Capabilities in AI: Complete Cross-Evaluation Results Using RFC-EVAL-001 Protocol</strong></em></p> <h3><em><strong>ABSTRACT </strong></em></h3> <p><em><strong>Background:</strong> As AI systems advance, there is increasing need for benchmarks that assess complex cognitive capabilities beyond narrow task performance. This study applies the RFC-EVAL-001 protocol—originally conceived within the Theoretical Potential Consciousness (TPC) framework—to evaluate multidimensional cognitive abilities in modern large language models (LLMs).</em></p> <p><em><strong>Methods:</strong> Six commercial LLMs (anonymized as IA1-IA6) were evaluated through 36 cross-evaluations. Each model assessed the others across four cognitive dimensions: Model Complexity (Cm), Temporal Horizon (Ht), Meta-Modeling (Mm), and Adaptive Flexibility (Fa). Inter-rater reliability was calculated using ICC(3,k), and cross-validation was performed by independent analysts.</em></p> <p><em><strong>Results:</strong> The protocol demonstrated excellent inter-rater reliability (average ICC(3,5) = 0.947). One system achieved NOVA-5 classification (γ = 0.932), while others scored at NOVA-4 level. Distinct cognitive profiles emerged: superior planning (Ht = 0.908), creative adaptation (Fa = 0.922), and balanced excellence. Epistemic calibration varied significantly across models. The complete dataset includes all 36 evaluations, Python analysis code, and protocol specifications.</em></p> <p><em><strong>Conclusions:</strong> This work provides a robust, replicable benchmark for comparing complex cognitive capabilities across AI systems. While explicitly not measuring consciousness, it offers a validated framework for assessing differential strengths in reasoning, planning, metacognition, and creativity, with practical tools for collaboration design (RFC-RC, RFC-CONEX).</em></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18684719
institution Zenodo
language spa
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle 30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
Bernal Díaz, Víctor Cristóbal
AI evaluation
cognitive benchmarking
large language models
inter-rater reliability
epistemic calibration
RFC-EVAL-001
cognitive interoperability
AI collaboration design
open dataset
evaluación de IA
benchmarking cognitivo
interoperabilidad cognitiva
diseño de colaboración IA
Teoría de la Potencialidad Consciente (TPC)
<h3><strong>RESUMEN EJECUTIVO:</strong></h3> <p>Este documento presenta la aplicación completa del protocolo RFC-EVAL-001 v1.1 para evaluar capacidades cognitivas complejas en seis sistemas de IA comerciales contemporáneos. Incluye 36 evaluaciones cruzadas ciegas, perfiles Ω dimensionales (Cm, Ht, Mm, Fa), clasificaciones NOVA, herramientas RFC-RC/RFC-CONEX para diseño de colaboraciones, y el dataset completo con código de replicación.</p> <p>Sobre la V2.1: Este documento ha sido sometido a un proceso de verificación metodológica independiente. Durante dicho proceso se identificaron y corrigieron errores en el cálculo de las puntuaciones Gamma (γ) y en el reporte de fiabilidad inter-evaluador presentes en una versión preliminar anterior.</p> <p>Sobre esta versión 2.2: Estado: Protocolo validado, datos verificados.</p> <p><em><strong>Benchmarking Complex Cognitive Capabilities in AI: Complete Cross-Evaluation Results Using RFC-EVAL-001 Protocol</strong></em></p> <h3><em><strong>ABSTRACT </strong></em></h3> <p><em><strong>Background:</strong> As AI systems advance, there is increasing need for benchmarks that assess complex cognitive capabilities beyond narrow task performance. This study applies the RFC-EVAL-001 protocol—originally conceived within the Theoretical Potential Consciousness (TPC) framework—to evaluate multidimensional cognitive abilities in modern large language models (LLMs).</em></p> <p><em><strong>Methods:</strong> Six commercial LLMs (anonymized as IA1-IA6) were evaluated through 36 cross-evaluations. Each model assessed the others across four cognitive dimensions: Model Complexity (Cm), Temporal Horizon (Ht), Meta-Modeling (Mm), and Adaptive Flexibility (Fa). Inter-rater reliability was calculated using ICC(3,k), and cross-validation was performed by independent analysts.</em></p> <p><em><strong>Results:</strong> The protocol demonstrated excellent inter-rater reliability (average ICC(3,5) = 0.947). One system achieved NOVA-5 classification (γ = 0.932), while others scored at NOVA-4 level. Distinct cognitive profiles emerged: superior planning (Ht = 0.908), creative adaptation (Fa = 0.922), and balanced excellence. Epistemic calibration varied significantly across models. The complete dataset includes all 36 evaluations, Python analysis code, and protocol specifications.</em></p> <p><em><strong>Conclusions:</strong> This work provides a robust, replicable benchmark for comparing complex cognitive capabilities across AI systems. While explicitly not measuring consciousness, it offers a validated framework for assessing differential strengths in reasoning, planning, metacognition, and creativity, with practical tools for collaboration design (RFC-RC, RFC-CONEX).</em></p>
title 30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
topic AI evaluation
cognitive benchmarking
large language models
inter-rater reliability
epistemic calibration
RFC-EVAL-001
cognitive interoperability
AI collaboration design
open dataset
evaluación de IA
benchmarking cognitivo
interoperabilidad cognitiva
diseño de colaboración IA
Teoría de la Potencialidad Consciente (TPC)
url https://doi.org/10.5281/zenodo.18684719