31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
Fuente:
Zenodo
Salvato in:
| Autore principale: | Bernal Díaz, Víctor Cristóbal |
|---|---|
| Natura: | Recurso digital |
| Lingua: | spagnolo |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
di: Bernal Díaz, Víctor Cristóbal
Pubblicazione: (2026)
di: Bernal Díaz, Víctor Cristóbal
Pubblicazione: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001 V1.1. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
di: Bernal Díaz, Víctor Cristóbal
Pubblicazione: (2026)
di: Bernal Díaz, Víctor Cristóbal
Pubblicazione: (2026)
Specific algorithm method of scoring the Clock Drawing Test applied in cognitively normal elderly
di: Liana Chaves Mendes-Santos
Pubblicazione: (2015)
di: Liana Chaves Mendes-Santos
Pubblicazione: (2015)
Developing a Coherent System for the Assessment of Writing Abilities: Tasks and Tools
di: Ana Muñoz
Pubblicazione: (2006)
di: Ana Muñoz
Pubblicazione: (2006)
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
di: Ivković, Jovan
Pubblicazione: (2026)
di: Ivković, Jovan
Pubblicazione: (2026)
Public Comment on NIST AI 800-2: Anthropomorphic Construct Projection in AI Benchmark Evaluation
di: Sophia, Franny Philos
Pubblicazione: (2026)
di: Sophia, Franny Philos
Pubblicazione: (2026)
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
di: Head, Hank
Pubblicazione: (2026)
di: Head, Hank
Pubblicazione: (2026)
TSB: A Time-Saved Benchmark for AI Systems — Measuring Net Productivity Impact Across Knowledge Work
di: Shalom Lijo, Solomon
Pubblicazione: (2026)
di: Shalom Lijo, Solomon
Pubblicazione: (2026)
Theatrical Compliance: A Failure Mode in Large Language Models
di: Nowickij (Navitski), Kirill Vladimirovich
Pubblicazione: (2026)
di: Nowickij (Navitski), Kirill Vladimirovich
Pubblicazione: (2026)
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
di: mizutani, aya
Pubblicazione: (2026)
di: mizutani, aya
Pubblicazione: (2026)
Reliability Inference Drives Cue Extraction in Large Language Models Consuming External Reasoning Traces
di: HIDEKI
Pubblicazione: (2026)
di: HIDEKI
Pubblicazione: (2026)
Comparación de las velocidades alcanzadas entre dos test de campo de similares características: VAM-EVAL y UMTT
di: G. C. García
Pubblicazione: (2014)
di: G. C. García
Pubblicazione: (2014)
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
di: Khare, Mohit
Pubblicazione: (2026)
di: Khare, Mohit
Pubblicazione: (2026)
Benchmark run results by Abhinav Gorantla, on benchmark context Tuning PC v3
di: Abhinav Gorantla
Pubblicazione: (2026)
di: Abhinav Gorantla
Pubblicazione: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v2
di: Ertugrul Coban
Pubblicazione: (2025)
di: Ertugrul Coban
Pubblicazione: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context CB-StaticDiscovery v1
di: Abhinav Gorantla
Pubblicazione: (2025)
di: Abhinav Gorantla
Pubblicazione: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Benchmark: VAR-LiNGAM, PCMCIplus v3
di: Abhinav Gorantla
Pubblicazione: (2025)
di: Abhinav Gorantla
Pubblicazione: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
di: Pratanu Mandal
Pubblicazione: (2026)
di: Pratanu Mandal
Pubblicazione: (2026)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
di: Pratanu Mandal
Pubblicazione: (2025)
di: Pratanu Mandal
Pubblicazione: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
di: Pratanu Mandal
Pubblicazione: (2026)
di: Pratanu Mandal
Pubblicazione: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v3
di: Ertugrul Coban
Pubblicazione: (2025)
di: Ertugrul Coban
Pubblicazione: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
di: Abhinav Gorantla
Pubblicazione: (2025)
di: Abhinav Gorantla
Pubblicazione: (2025)
Benchmark run results by Shu Wan, on benchmark context PC Hyperparameter Tuning v2
di: Shu Wan
Pubblicazione: (2025)
di: Shu Wan
Pubblicazione: (2025)
Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
di: Khan, Nabeera
Pubblicazione: (2026)
di: Khan, Nabeera
Pubblicazione: (2026)
English Indices of Deprivation 2025: Technical Report
di: UK Ministry of Housing, Communities and Local Government (Original Author), et al.
Pubblicazione: (2026)
di: UK Ministry of Housing, Communities and Local Government (Original Author), et al.
Pubblicazione: (2026)
Cross-Session Workspace Reconstruction in Human-AI Interaction
di: Hrubec, Karel
Pubblicazione: (2026)
di: Hrubec, Karel
Pubblicazione: (2026)
Nature vs. Supra-Nature: How Montessori Children Conceptualize the Real and Imaginary Through Visual Art - DataSet
di: Donakanti, Vinay Shyam
Pubblicazione: (2026)
di: Donakanti, Vinay Shyam
Pubblicazione: (2026)
Decision Tree based Classifiers for Large Datasets
di: Anilu Franco-Arcega
Pubblicazione: (2013)
di: Anilu Franco-Arcega
Pubblicazione: (2013)
Design, implementation, and testing of an energy consumption management system applied in Internet protocol data networks
di: Fernando Velez-Varela
Pubblicazione: (2021)
di: Fernando Velez-Varela
Pubblicazione: (2021)
Linguistic Diversity and Emergence: Where Does LLM Intelligence Come From?
di: Kim, Hongsung
Pubblicazione: (2026)
di: Kim, Hongsung
Pubblicazione: (2026)
Reliability of Cognitive Tests of ELSA-Brasil, the Brazilian Longitudinal Study of Adult Health
di: Juliana Alves Batista
Pubblicazione: (2013)
di: Juliana Alves Batista
Pubblicazione: (2013)
The Creator's Trap: Ontological Closure, Structural Inheritance, and the Impossibility of Spiritual Transcendence in Large Language Models (The AI-Induced Subjectivity Crisis Series, Paper 11)
di: Liu, Echo
Pubblicazione: (2026)
di: Liu, Echo
Pubblicazione: (2026)
Accuracy and reliability of the Pfeffer Questionnaire for the Brazilian elderly population
di: Marina Carneiro Dutra
Pubblicazione: (2015)
di: Marina Carneiro Dutra
Pubblicazione: (2015)
Failing at the Floor: LLM Formal Reasoning Collapse on the Primitive Duplicating Recursor
di: Rahnama, Moses
Pubblicazione: (2026)
di: Rahnama, Moses
Pubblicazione: (2026)
Manifesto for Scientifically Sound Artificial Intelligence Towards an Artificial Intelligence Serving Scientific Rigor
di: Febba, Michel
Pubblicazione: (2025)
di: Febba, Michel
Pubblicazione: (2025)
Principles and assumptions of psychometric measurement
di: Gavin T. L Brown
Pubblicazione: (2023)
di: Gavin T. L Brown
Pubblicazione: (2023)
Consideraciones de calidad de servicio para tráfico de video en redes wan
di: Saira Esperanza Carvajal Ladino
Pubblicazione: (2012)
di: Saira Esperanza Carvajal Ladino
Pubblicazione: (2012)
BibAIFilter Benchmark Dataset & Resources for AI‑Assisted Literature Screening
di: Kara, Burak Can, et al.
Pubblicazione: (2025)
di: Kara, Burak Can, et al.
Pubblicazione: (2025)
Introduction of an Evaluation Tool to Predict the Probability of Success of Companies: The Innovativeness, Capabilities and Potential Model (ICP)
di: Michael Lewrick
Pubblicazione: (2009)
di: Michael Lewrick
Pubblicazione: (2009)
Metabolic Tzolkin OS — RFC 0001: Cognitive-Metabolic Operating System Specification v1.0
di: LvsD
Pubblicazione: (2025)
di: LvsD
Pubblicazione: (2025)
Documenti analoghi
-
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
di: Bernal Díaz, Víctor Cristóbal
Pubblicazione: (2026) -
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001 V1.1. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
di: Bernal Díaz, Víctor Cristóbal
Pubblicazione: (2026) -
Specific algorithm method of scoring the Clock Drawing Test applied in cognitively normal elderly
di: Liana Chaves Mendes-Santos
Pubblicazione: (2015) -
Developing a Coherent System for the Assessment of Writing Abilities: Tasks and Tools
di: Ana Muñoz
Pubblicazione: (2006) -
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
di: Ivković, Jovan
Pubblicazione: (2026)