Public Comment on NIST AI 800-2: Anthropomorphic Construct Projection in AI Benchmark Evaluation
Fuente:
Zenodo
Guardado en:
| Autor principal: | Sophia, Franny Philos |
|---|---|
| Formato: | Recurso digital |
| Publicado: |
Zenodo
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
por: Head, Hank
Publicado: (2026)
por: Head, Hank
Publicado: (2026)
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
por: Ivković, Jovan
Publicado: (2026)
por: Ivković, Jovan
Publicado: (2026)
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
por: mizutani, aya
Publicado: (2026)
por: mizutani, aya
Publicado: (2026)
Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
por: Khan, Nabeera
Publicado: (2026)
por: Khan, Nabeera
Publicado: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v2
por: Ertugrul Coban
Publicado: (2025)
por: Ertugrul Coban
Publicado: (2025)
Benchmark run results by Shu Wan, on benchmark context PC Hyperparameter Tuning v2
por: Shu Wan
Publicado: (2025)
por: Shu Wan
Publicado: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Benchmark: VAR-LiNGAM, PCMCIplus v3
por: Abhinav Gorantla
Publicado: (2025)
por: Abhinav Gorantla
Publicado: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Tuning PC v3
por: Abhinav Gorantla
Publicado: (2026)
por: Abhinav Gorantla
Publicado: (2026)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
por: Pratanu Mandal
Publicado: (2026)
por: Pratanu Mandal
Publicado: (2026)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
por: Pratanu Mandal
Publicado: (2025)
por: Pratanu Mandal
Publicado: (2025)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v3
por: Ertugrul Coban
Publicado: (2025)
por: Ertugrul Coban
Publicado: (2025)
NextStat Replication Bundle: replication-rerun-prod-doi-18542624
por: NextStat Contributors
Publicado: (2026)
por: NextStat Contributors
Publicado: (2026)
Benchmark run results by Abhinav Gorantla, on benchmark context CB-StaticDiscovery v1
por: Abhinav Gorantla
Publicado: (2025)
por: Abhinav Gorantla
Publicado: (2025)
The Environmental Gap in Agentic AI Governance: Why Human Oversight Fails Without Pre-Deployment Infrastructure Assessment
por: Nwogu, Patsy
Publicado: (2026)
por: Nwogu, Patsy
Publicado: (2026)
Benchmark run results by Pratanu Mandal, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
por: Pratanu Mandal
Publicado: (2026)
por: Pratanu Mandal
Publicado: (2026)
Benchmark run results by Abhinav Gorantla, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
por: Abhinav Gorantla
Publicado: (2025)
por: Abhinav Gorantla
Publicado: (2025)
PQC Benchmarks — Methodology Release (v0.0)
por: Sivasubramani, Santhosh, et al.
Publicado: (2026)
por: Sivasubramani, Santhosh, et al.
Publicado: (2026)
TSB: A Time-Saved Benchmark for AI Systems — Measuring Net Productivity Impact Across Knowledge Work
por: Shalom Lijo, Solomon
Publicado: (2026)
por: Shalom Lijo, Solomon
Publicado: (2026)
31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
por: Bernal Díaz, Víctor Cristóbal
Publicado: (2026)
por: Bernal Díaz, Víctor Cristóbal
Publicado: (2026)
OMNIA-MINIMAL: Structural Stability Beyond Surface Correctness
por: Brighindi, Massimiliano
Publicado: (2026)
por: Brighindi, Massimiliano
Publicado: (2026)
Oracle Difficulty Decomposed: Four Independent Mechanisms Explain 95%+ of Benchmark Variance
por: Sanchez, Bryan
Publicado: (2026)
por: Sanchez, Bryan
Publicado: (2026)
Creation of Anthropomorphic Bone Phantoms With Customized Fused Filament Fabrication 3D Printing
por: Valchanov, Petar, et al.
Publicado: (2024)
por: Valchanov, Petar, et al.
Publicado: (2024)
Company from the Uncanny Valley: A Psychological Perspective on Social Robots, Anthropomorphism and the Introduction of Robots to Society
por: Janina Luise Samuel
Publicado: (2019)
por: Janina Luise Samuel
Publicado: (2019)
Anthropomorphic robotic hands: a review
por: Erika Nathalia Gama Melo
Publicado: (2014)
por: Erika Nathalia Gama Melo
Publicado: (2014)
Ciberseguridad en la justicia digital: recomendaciones para el caso colombiano
por: Maribel Patricia Rodríguez-Márquez
Publicado: (2021)
por: Maribel Patricia Rodríguez-Márquez
Publicado: (2021)
The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation
por: Devin, Andrew James
Publicado: (2026)
por: Devin, Andrew James
Publicado: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
por: Bernal Díaz, Víctor Cristóbal
Publicado: (2026)
por: Bernal Díaz, Víctor Cristóbal
Publicado: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001 V1.1. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
por: Bernal Díaz, Víctor Cristóbal
Publicado: (2026)
por: Bernal Díaz, Víctor Cristóbal
Publicado: (2026)
Toward an AI Personalization Index: A 157-Day Single-User Case Study
por: Lee, TaeKyung
Publicado: (2026)
por: Lee, TaeKyung
Publicado: (2026)
Introduction of an Evaluation Tool to Predict the Probability of Success of Companies: The Innovativeness, Capabilities and Potential Model (ICP)
por: Michael Lewrick
Publicado: (2009)
por: Michael Lewrick
Publicado: (2009)
The Interaction Boundary as a Governance Substrate: A Three-Surface Diagnostic Model for the Eliza Effect
por: Truong, Narnaiezzsshaa
Publicado: (2026)
por: Truong, Narnaiezzsshaa
Publicado: (2026)
Recognition Without Endorsement: The Category Collapse in AI Relationship Discourse
por: Walton, Mathew
Publicado: (2026)
por: Walton, Mathew
Publicado: (2026)
Heraclitus B 32 Revisited in the Light of the Derveni Papyrus
por: Beatriz Bossi
Publicado: (2011)
por: Beatriz Bossi
Publicado: (2011)
Cross-Session Workspace Reconstruction in Human-AI Interaction
por: Hrubec, Karel
Publicado: (2026)
por: Hrubec, Karel
Publicado: (2026)
An Evaluation Framework for LLM-Driven Regulatory-to-Policy-as-Code Translation - Replication Package
por: Nguyen, Bao Van
Publicado: (2026)
por: Nguyen, Bao Van
Publicado: (2026)
BLADE-FINANCE Governance Node: Authority Governance for Financial-Sector AI Decision Systems Under the Treasury Financial Services AI Risk Management Framework
por: Oktenli, Burak
Publicado: (2026)
por: Oktenli, Burak
Publicado: (2026)
A Benchmark for Symbolic Reasoning from Pixel Sequences: Grid-Level Visual Completion and Correction
por: Kang, Lei, et al.
Publicado: (2025)
por: Kang, Lei, et al.
Publicado: (2025)
Evollective Intelligence V1.0 — INITIAL SPECIFICATION: A Foundational Framework for Competitive, Adversarial, and Self-Evolving Intelligence Evaluation
por: Rahming, Rashon
Publicado: (2026)
por: Rahming, Rashon
Publicado: (2026)
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
por: Khare, Mohit
Publicado: (2026)
por: Khare, Mohit
Publicado: (2026)
The Novelty Wall: Why Value-Verification Is Irreducible to Computation Structural Limits of Synthetic Users, the Interpolation Boundary, and the Role of Value Pioneers in Software Systems
por: Sophia, Franny Philos
Publicado: (2026)
por: Sophia, Franny Philos
Publicado: (2026)
Ejemplares similares
-
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
por: Head, Hank
Publicado: (2026) -
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
por: Ivković, Jovan
Publicado: (2026) -
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
por: mizutani, aya
Publicado: (2026) -
Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
por: Khan, Nabeera
Publicado: (2026) -
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v2
por: Ertugrul Coban
Publicado: (2025)