30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | |
|---|---|
| Format: | Recurso digital |
| Sprache: | Spanisch |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901750034726912 |
|---|---|
| author | Bernal Díaz, Víctor Cristóbal |
| author_facet | Bernal Díaz, Víctor Cristóbal |
| contents | <h3><strong>RESUMEN EJECUTIVO:</strong></h3> <p>Este documento presenta la aplicación completa del protocolo RFC-EVAL-001 v1.1 para evaluar capacidades cognitivas complejas en seis sistemas de IA comerciales contemporáneos. Incluye 36 evaluaciones cruzadas ciegas, perfiles Ω dimensionales (Cm, Ht, Mm, Fa), clasificaciones NOVA, herramientas RFC-RC/RFC-CONEX para diseño de colaboraciones, y el dataset completo con código de replicación.</p> <p>Sobre la V2.1: Este documento ha sido sometido a un proceso de verificación metodológica independiente. Durante dicho proceso se identificaron y corrigieron errores en el cálculo de las puntuaciones Gamma (γ) y en el reporte de fiabilidad inter-evaluador presentes en una versión preliminar anterior.</p> <p>Sobre esta versión 2.2: Estado: Protocolo validado, datos verificados.</p> <p><em><strong>Benchmarking Complex Cognitive Capabilities in AI: Complete Cross-Evaluation Results Using RFC-EVAL-001 Protocol</strong></em></p> <h3><em><strong>ABSTRACT </strong></em></h3> <p><em><strong>Background:</strong> As AI systems advance, there is increasing need for benchmarks that assess complex cognitive capabilities beyond narrow task performance. This study applies the RFC-EVAL-001 protocol—originally conceived within the Theoretical Potential Consciousness (TPC) framework—to evaluate multidimensional cognitive abilities in modern large language models (LLMs).</em></p> <p><em><strong>Methods:</strong> Six commercial LLMs (anonymized as IA1-IA6) were evaluated through 36 cross-evaluations. Each model assessed the others across four cognitive dimensions: Model Complexity (Cm), Temporal Horizon (Ht), Meta-Modeling (Mm), and Adaptive Flexibility (Fa). Inter-rater reliability was calculated using ICC(3,k), and cross-validation was performed by independent analysts.</em></p> <p><em><strong>Results:</strong> The protocol demonstrated excellent inter-rater reliability (average ICC(3,5) = 0.947). One system achieved NOVA-5 classification (γ = 0.932), while others scored at NOVA-4 level. Distinct cognitive profiles emerged: superior planning (Ht = 0.908), creative adaptation (Fa = 0.922), and balanced excellence. Epistemic calibration varied significantly across models. The complete dataset includes all 36 evaluations, Python analysis code, and protocol specifications.</em></p> <p><em><strong>Conclusions:</strong> This work provides a robust, replicable benchmark for comparing complex cognitive capabilities across AI systems. While explicitly not measuring consciousness, it offers a validated framework for assessing differential strengths in reasoning, planning, metacognition, and creativity, with practical tools for collaboration design (RFC-RC, RFC-CONEX).</em></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18684719 |
| institution | Zenodo |
| language | spa |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | 30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES. Bernal Díaz, Víctor Cristóbal AI evaluation cognitive benchmarking large language models inter-rater reliability epistemic calibration RFC-EVAL-001 cognitive interoperability AI collaboration design open dataset evaluación de IA benchmarking cognitivo interoperabilidad cognitiva diseño de colaboración IA Teoría de la Potencialidad Consciente (TPC) <h3><strong>RESUMEN EJECUTIVO:</strong></h3> <p>Este documento presenta la aplicación completa del protocolo RFC-EVAL-001 v1.1 para evaluar capacidades cognitivas complejas en seis sistemas de IA comerciales contemporáneos. Incluye 36 evaluaciones cruzadas ciegas, perfiles Ω dimensionales (Cm, Ht, Mm, Fa), clasificaciones NOVA, herramientas RFC-RC/RFC-CONEX para diseño de colaboraciones, y el dataset completo con código de replicación.</p> <p>Sobre la V2.1: Este documento ha sido sometido a un proceso de verificación metodológica independiente. Durante dicho proceso se identificaron y corrigieron errores en el cálculo de las puntuaciones Gamma (γ) y en el reporte de fiabilidad inter-evaluador presentes en una versión preliminar anterior.</p> <p>Sobre esta versión 2.2: Estado: Protocolo validado, datos verificados.</p> <p><em><strong>Benchmarking Complex Cognitive Capabilities in AI: Complete Cross-Evaluation Results Using RFC-EVAL-001 Protocol</strong></em></p> <h3><em><strong>ABSTRACT </strong></em></h3> <p><em><strong>Background:</strong> As AI systems advance, there is increasing need for benchmarks that assess complex cognitive capabilities beyond narrow task performance. This study applies the RFC-EVAL-001 protocol—originally conceived within the Theoretical Potential Consciousness (TPC) framework—to evaluate multidimensional cognitive abilities in modern large language models (LLMs).</em></p> <p><em><strong>Methods:</strong> Six commercial LLMs (anonymized as IA1-IA6) were evaluated through 36 cross-evaluations. Each model assessed the others across four cognitive dimensions: Model Complexity (Cm), Temporal Horizon (Ht), Meta-Modeling (Mm), and Adaptive Flexibility (Fa). Inter-rater reliability was calculated using ICC(3,k), and cross-validation was performed by independent analysts.</em></p> <p><em><strong>Results:</strong> The protocol demonstrated excellent inter-rater reliability (average ICC(3,5) = 0.947). One system achieved NOVA-5 classification (γ = 0.932), while others scored at NOVA-4 level. Distinct cognitive profiles emerged: superior planning (Ht = 0.908), creative adaptation (Fa = 0.922), and balanced excellence. Epistemic calibration varied significantly across models. The complete dataset includes all 36 evaluations, Python analysis code, and protocol specifications.</em></p> <p><em><strong>Conclusions:</strong> This work provides a robust, replicable benchmark for comparing complex cognitive capabilities across AI systems. While explicitly not measuring consciousness, it offers a validated framework for assessing differential strengths in reasoning, planning, metacognition, and creativity, with practical tools for collaboration design (RFC-RC, RFC-CONEX).</em></p> |
| title | 30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES. |
| topic | AI evaluation cognitive benchmarking large language models inter-rater reliability epistemic calibration RFC-EVAL-001 cognitive interoperability AI collaboration design open dataset evaluación de IA benchmarking cognitivo interoperabilidad cognitiva diseño de colaboración IA Teoría de la Potencialidad Consciente (TPC) |
| url | https://doi.org/10.5281/zenodo.18684719 |