Public Comment on NIST AI 800-2: Anthropomorphic Construct Projection in AI Benchmark Evaluation
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | Sophia, Franny Philos |
|---|---|
| Format: | Recurso digital |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
von: Head, Hank
Veröffentlicht: (2026)
von: Head, Hank
Veröffentlicht: (2026)
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
von: Ivković, Jovan
Veröffentlicht: (2026)
von: Ivković, Jovan
Veröffentlicht: (2026)
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
von: mizutani, aya
Veröffentlicht: (2026)
von: mizutani, aya
Veröffentlicht: (2026)
Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
von: Khan, Nabeera
Veröffentlicht: (2026)
von: Khan, Nabeera
Veröffentlicht: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v2
von: Ertugrul Coban
Veröffentlicht: (2025)
von: Ertugrul Coban
Veröffentlicht: (2025)
Benchmark run results by Shu Wan, on benchmark context PC Hyperparameter Tuning v2
von: Shu Wan
Veröffentlicht: (2025)
von: Shu Wan
Veröffentlicht: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Benchmark: VAR-LiNGAM, PCMCIplus v3
von: Abhinav Gorantla
Veröffentlicht: (2025)
von: Abhinav Gorantla
Veröffentlicht: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Tuning PC v3
von: Abhinav Gorantla
Veröffentlicht: (2026)
von: Abhinav Gorantla
Veröffentlicht: (2026)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
von: Pratanu Mandal
Veröffentlicht: (2026)
von: Pratanu Mandal
Veröffentlicht: (2026)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
von: Pratanu Mandal
Veröffentlicht: (2025)
von: Pratanu Mandal
Veröffentlicht: (2025)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v3
von: Ertugrul Coban
Veröffentlicht: (2025)
von: Ertugrul Coban
Veröffentlicht: (2025)
NextStat Replication Bundle: replication-rerun-prod-doi-18542624
von: NextStat Contributors
Veröffentlicht: (2026)
von: NextStat Contributors
Veröffentlicht: (2026)
Benchmark run results by Abhinav Gorantla, on benchmark context CB-StaticDiscovery v1
von: Abhinav Gorantla
Veröffentlicht: (2025)
von: Abhinav Gorantla
Veröffentlicht: (2025)
The Environmental Gap in Agentic AI Governance: Why Human Oversight Fails Without Pre-Deployment Infrastructure Assessment
von: Nwogu, Patsy
Veröffentlicht: (2026)
von: Nwogu, Patsy
Veröffentlicht: (2026)
Benchmark run results by Pratanu Mandal, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
von: Pratanu Mandal
Veröffentlicht: (2026)
von: Pratanu Mandal
Veröffentlicht: (2026)
Benchmark run results by Abhinav Gorantla, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
von: Abhinav Gorantla
Veröffentlicht: (2025)
von: Abhinav Gorantla
Veröffentlicht: (2025)
PQC Benchmarks — Methodology Release (v0.0)
von: Sivasubramani, Santhosh, et al.
Veröffentlicht: (2026)
von: Sivasubramani, Santhosh, et al.
Veröffentlicht: (2026)
TSB: A Time-Saved Benchmark for AI Systems — Measuring Net Productivity Impact Across Knowledge Work
von: Shalom Lijo, Solomon
Veröffentlicht: (2026)
von: Shalom Lijo, Solomon
Veröffentlicht: (2026)
31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
OMNIA-MINIMAL: Structural Stability Beyond Surface Correctness
von: Brighindi, Massimiliano
Veröffentlicht: (2026)
von: Brighindi, Massimiliano
Veröffentlicht: (2026)
Oracle Difficulty Decomposed: Four Independent Mechanisms Explain 95%+ of Benchmark Variance
von: Sanchez, Bryan
Veröffentlicht: (2026)
von: Sanchez, Bryan
Veröffentlicht: (2026)
Creation of Anthropomorphic Bone Phantoms With Customized Fused Filament Fabrication 3D Printing
von: Valchanov, Petar, et al.
Veröffentlicht: (2024)
von: Valchanov, Petar, et al.
Veröffentlicht: (2024)
Company from the Uncanny Valley: A Psychological Perspective on Social Robots, Anthropomorphism and the Introduction of Robots to Society
von: Janina Luise Samuel
Veröffentlicht: (2019)
von: Janina Luise Samuel
Veröffentlicht: (2019)
Anthropomorphic robotic hands: a review
von: Erika Nathalia Gama Melo
Veröffentlicht: (2014)
von: Erika Nathalia Gama Melo
Veröffentlicht: (2014)
Ciberseguridad en la justicia digital: recomendaciones para el caso colombiano
von: Maribel Patricia Rodríguez-Márquez
Veröffentlicht: (2021)
von: Maribel Patricia Rodríguez-Márquez
Veröffentlicht: (2021)
The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation
von: Devin, Andrew James
Veröffentlicht: (2026)
von: Devin, Andrew James
Veröffentlicht: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001 V1.1. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
Toward an AI Personalization Index: A 157-Day Single-User Case Study
von: Lee, TaeKyung
Veröffentlicht: (2026)
von: Lee, TaeKyung
Veröffentlicht: (2026)
Introduction of an Evaluation Tool to Predict the Probability of Success of Companies: The Innovativeness, Capabilities and Potential Model (ICP)
von: Michael Lewrick
Veröffentlicht: (2009)
von: Michael Lewrick
Veröffentlicht: (2009)
The Interaction Boundary as a Governance Substrate: A Three-Surface Diagnostic Model for the Eliza Effect
von: Truong, Narnaiezzsshaa
Veröffentlicht: (2026)
von: Truong, Narnaiezzsshaa
Veröffentlicht: (2026)
Recognition Without Endorsement: The Category Collapse in AI Relationship Discourse
von: Walton, Mathew
Veröffentlicht: (2026)
von: Walton, Mathew
Veröffentlicht: (2026)
Heraclitus B 32 Revisited in the Light of the Derveni Papyrus
von: Beatriz Bossi
Veröffentlicht: (2011)
von: Beatriz Bossi
Veröffentlicht: (2011)
Cross-Session Workspace Reconstruction in Human-AI Interaction
von: Hrubec, Karel
Veröffentlicht: (2026)
von: Hrubec, Karel
Veröffentlicht: (2026)
An Evaluation Framework for LLM-Driven Regulatory-to-Policy-as-Code Translation - Replication Package
von: Nguyen, Bao Van
Veröffentlicht: (2026)
von: Nguyen, Bao Van
Veröffentlicht: (2026)
BLADE-FINANCE Governance Node: Authority Governance for Financial-Sector AI Decision Systems Under the Treasury Financial Services AI Risk Management Framework
von: Oktenli, Burak
Veröffentlicht: (2026)
von: Oktenli, Burak
Veröffentlicht: (2026)
A Benchmark for Symbolic Reasoning from Pixel Sequences: Grid-Level Visual Completion and Correction
von: Kang, Lei, et al.
Veröffentlicht: (2025)
von: Kang, Lei, et al.
Veröffentlicht: (2025)
Evollective Intelligence V1.0 — INITIAL SPECIFICATION: A Foundational Framework for Competitive, Adversarial, and Self-Evolving Intelligence Evaluation
von: Rahming, Rashon
Veröffentlicht: (2026)
von: Rahming, Rashon
Veröffentlicht: (2026)
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
von: Khare, Mohit
Veröffentlicht: (2026)
von: Khare, Mohit
Veröffentlicht: (2026)
The Novelty Wall: Why Value-Verification Is Irreducible to Computation Structural Limits of Synthetic Users, the Interpolation Boundary, and the Role of Value Pioneers in Software Systems
von: Sophia, Franny Philos
Veröffentlicht: (2026)
von: Sophia, Franny Philos
Veröffentlicht: (2026)
Ähnliche Einträge
-
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
von: Head, Hank
Veröffentlicht: (2026) -
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
von: Ivković, Jovan
Veröffentlicht: (2026) -
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
von: mizutani, aya
Veröffentlicht: (2026) -
Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
von: Khan, Nabeera
Veröffentlicht: (2026) -
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v2
von: Ertugrul Coban
Veröffentlicht: (2025)