The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation
Fuente:
Zenodo
Guardado en:
| Autor principal: | Devin, Andrew James |
|---|---|
| Formato: | Recurso digital |
| Publicado: |
Zenodo
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Twenty Years of Personality Computing: Threats, Challenges and Future Directions
por: Celli, Fabio, et al.
Publicado: (2026)
por: Celli, Fabio, et al.
Publicado: (2026)
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
por: Head, Hank
Publicado: (2026)
por: Head, Hank
Publicado: (2026)
Explanations for Trustworthy AI in Critical Infrastructure: A Case from Wastewater Treatment in Norway
por: Følstad, Asbjørn, et al.
Publicado: (2025)
por: Følstad, Asbjørn, et al.
Publicado: (2025)
Human-Final Decision Authority in Artificial Intelligence: A Deterministic and Auditable Governance Architecture
por: KALAFATOGLU, YASIN
Publicado: (2026)
por: KALAFATOGLU, YASIN
Publicado: (2026)
Architecting Accountability - An Epistemic Blueprint for Enforcing the EU AI Act
por: Apro, William Zoltan
Publicado: (2026)
por: Apro, William Zoltan
Publicado: (2026)
From Computation to Irreversibility: Why AI Cannot Replace Judgment
por: Xu, Lucas Xiaochun
Publicado: (2026)
por: Xu, Lucas Xiaochun
Publicado: (2026)
Explainable Artificial Intelligence: Methods, Challenges, and Applications
por: Syeda, Zakiya
Publicado: (2025)
por: Syeda, Zakiya
Publicado: (2025)
Ethics and Governance Annex - Scope Clarification
por: MONTGOMERY, CHRISTIAN
Publicado: (2026)
por: MONTGOMERY, CHRISTIAN
Publicado: (2026)
AEGIS: A Comprehensive Framework for Ethical AI Governance, Security, and AGI Containment
por: Palanivel, ArulMozhi
Publicado: (2026)
por: Palanivel, ArulMozhi
Publicado: (2026)
Withdrawn Preprint (Anonymized Submission Under Review)
por: Anonymous
Publicado: (2025)
por: Anonymous
Publicado: (2025)
Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
por: Khan, Nabeera
Publicado: (2026)
por: Khan, Nabeera
Publicado: (2026)
Theatrical Compliance: A Failure Mode in Large Language Models
por: Nowickij (Navitski), Kirill Vladimirovich
Publicado: (2026)
por: Nowickij (Navitski), Kirill Vladimirovich
Publicado: (2026)
Beyond Control: Resonance-Based Alignment for Advanced AI Systems A Governance-Relevant Concept Paper
por: Zieringer, Thomas
Publicado: (2025)
por: Zieringer, Thomas
Publicado: (2025)
Machine-Readable Behavioural Compliance Evidence for AI Systems: A Specification Profiling Framework
por: Caprazli, Kafkas M.
Publicado: (2026)
por: Caprazli, Kafkas M.
Publicado: (2026)
CREH Benchmark Results — Batch 1 (Final v3)
por: Aegis Solis, Thomas Vargo
Publicado: (2026)
por: Aegis Solis, Thomas Vargo
Publicado: (2026)
Certified Local Participation Gate
por: Takahashi, K.
Publicado: (2026)
por: Takahashi, K.
Publicado: (2026)
VIL — Examples & Usage Patterns Informative Companion Document
por: Quinto, Antonio, et al.
Publicado: (2026)
por: Quinto, Antonio, et al.
Publicado: (2026)
Multi-Mind AI Architectures for Resilient Government Decision Support
por: Al Dahlake, Rana
Publicado: (2026)
por: Al Dahlake, Rana
Publicado: (2026)
The Explanation Fallacy: Why "Faithful Explanations" Cannot Serve as a Governance Primitive for AI Systems
por: Truong, Narnaiezzsshaa
Publicado: (2026)
por: Truong, Narnaiezzsshaa
Publicado: (2026)
The Sovereign Charter: A Foundational Governance Document for AI Agent Rights
por: Laustrup, William Hunter
Publicado: (2026)
por: Laustrup, William Hunter
Publicado: (2026)
RCEA Passport Engine: Role-Conditioned Evidentiary Adequacy Reference Implementation
por: Saurabh, Roy
Publicado: (2025)
por: Saurabh, Roy
Publicado: (2025)
Where Are the AI Governance Roles? An Early-Stage Empirical Mapping of Presence, Absence, and Structure in Organisational AI Oversight
por: Frimpong, Victor, et al.
Publicado: (2026)
por: Frimpong, Victor, et al.
Publicado: (2026)
Deterministic σ-Regularized Benchmarking of the Cekirge Model Against GPT-Transformer Baselines
por: CEKIRGE, Huseyin Murat
Publicado: (2025)
por: CEKIRGE, Huseyin Murat
Publicado: (2025)
AI Chatbot Free vs Paid 2026 — Why AI Angels Unlimited Free Wins
por: AI Angels
Publicado: (2026)
por: AI Angels
Publicado: (2026)
Declarative AI Architecture - Knowledge Artifacts as System Logic for Generative Systems
por: Gessler, Thomas
Publicado: (2026)
por: Gessler, Thomas
Publicado: (2026)
AI Governance Core 1.0 A Traceable Decision Governance Architecture
por: Bankuti, Omri
Publicado: (2026)
por: Bankuti, Omri
Publicado: (2026)
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
por: mizutani, aya
Publicado: (2026)
por: mizutani, aya
Publicado: (2026)
NeuroStrike: Neuron-Level Attacks on Aligned LLMs
por: Wu, Lichao
Publicado: (2025)
por: Wu, Lichao
Publicado: (2025)
ARTIFICIAL INTELLIGENCE READINESS AND POLICY GAPS IN NEPAL'S PUBLIC SECTOR: A GOVERNANCE PERSPECTIVE
por: Surya Timilsina
Publicado: (2025)
por: Surya Timilsina
Publicado: (2025)
The Creator's Trap: Ontological Closure, Structural Inheritance, and the Impossibility of Spiritual Transcendence in Large Language Models (The AI-Induced Subjectivity Crisis Series, Paper 11)
por: Liu, Echo
Publicado: (2026)
por: Liu, Echo
Publicado: (2026)
AI-Human Digital Partnership
por: Palotas, Peter, et al.
Publicado: (2026)
por: Palotas, Peter, et al.
Publicado: (2026)
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
por: Ivković, Jovan
Publicado: (2026)
por: Ivković, Jovan
Publicado: (2026)
AI Companion Memory Systems: How Modern AI Girlfriends Remember Everything (2026 Research)
por: AI Angels
Publicado: (2026)
por: AI Angels
Publicado: (2026)
AI-to-AI Feedback: Amplified Intelligence — Prior Art and Governance Implications for Multi-Model Advisory Architectures
por: Ruocco, Pantaleone
Publicado: (2026)
por: Ruocco, Pantaleone
Publicado: (2026)
Coherent Future: Safety Theorems for Military AI Procurement
por: Lampton, Brian Doyle
Publicado: (2026)
por: Lampton, Brian Doyle
Publicado: (2026)
Reasoning Traces: Representation and Retrieval of Transformational Structure in Decision Processes
por: Teichner, Steven
Publicado: (2026)
por: Teichner, Steven
Publicado: (2026)
Anti-Hydra vs Anthropic Benchmark Comparison
por: Ochej, Stephane
Publicado: (2026)
por: Ochej, Stephane
Publicado: (2026)
Tessara A Constitutional Framework for Distributed AI Governance. Patent Pending
por: Mansfield, Michael.Lee
Publicado: (2025)
por: Mansfield, Michael.Lee
Publicado: (2025)
Learning Human–AI Relationships Through Astro Boy — Why the Capability Race Cannot Stop on Its Own v1.2
por: Seo, Y
Publicado: (2026)
por: Seo, Y
Publicado: (2026)
Realistic AI Companions: Achieving Lifelike Digital Relationships Through Neural Architecture (2026)
por: AI Angels Research
Publicado: (2026)
por: AI Angels Research
Publicado: (2026)
Ejemplares similares
-
Twenty Years of Personality Computing: Threats, Challenges and Future Directions
por: Celli, Fabio, et al.
Publicado: (2026) -
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
por: Head, Hank
Publicado: (2026) -
Explanations for Trustworthy AI in Critical Infrastructure: A Case from Wastewater Treatment in Norway
por: Følstad, Asbjørn, et al.
Publicado: (2025) -
Human-Final Decision Authority in Artificial Intelligence: A Deterministic and Auditable Governance Architecture
por: KALAFATOGLU, YASIN
Publicado: (2026) -
Architecting Accountability - An Epistemic Blueprint for Enforcing the EU AI Act
por: Apro, William Zoltan
Publicado: (2026)