Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
Fuente:
Zenodo
Saved in:
| Main Author: | Khan, Nabeera |
|---|---|
| Format: | Recurso digital |
| Language: | English |
| Published: |
Zenodo
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
by: Khare, Mohit
Published: (2026)
by: Khare, Mohit
Published: (2026)
Gemini Update Clinical decision support based on Bevacizumab cancer trials and pushing the limitations of advanced LLMs
by: Kawchak, Kevin
Published: (2025)
by: Kawchak, Kevin
Published: (2025)
The Brain Problem: Creative Constraint Optimization in Large Language Models
by: Marinello, Nicola, et al.
Published: (2026)
by: Marinello, Nicola, et al.
Published: (2026)
Estimating the Impact of Automation on Vocational Education: The Case of Technical Courses
by: Lima, Yuri, et al.
Published: (2024)
by: Lima, Yuri, et al.
Published: (2024)
LACF Emotional Paradigm: A Personalized Artificial Nervous System for Human-AI Alignment
by: Ochej, Stephane, et al.
Published: (2026)
by: Ochej, Stephane, et al.
Published: (2026)
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
by: Head, Hank
Published: (2026)
by: Head, Hank
Published: (2026)
Public Comment on NIST AI 800-2: Anthropomorphic Construct Projection in AI Benchmark Evaluation
by: Sophia, Franny Philos
Published: (2026)
by: Sophia, Franny Philos
Published: (2026)
When AI Tells You What You Want to Hear: Sycophantic Behavior of Large Language Models in Dementia Care Settings
by: Kolb, Christian
Published: (2026)
by: Kolb, Christian
Published: (2026)
Anima AI Community Pulse Dataset
by: AI Companion Picker, et al.
Published: (2026)
by: AI Companion Picker, et al.
Published: (2026)
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
by: Ivković, Jovan
Published: (2026)
by: Ivković, Jovan
Published: (2026)
Theatrical Compliance: A Failure Mode in Large Language Models
by: Nowickij (Navitski), Kirill Vladimirovich
Published: (2026)
by: Nowickij (Navitski), Kirill Vladimirovich
Published: (2026)
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
by: mizutani, aya
Published: (2026)
by: mizutani, aya
Published: (2026)
Persona, Shadow, and Cheap Coherence: A Jungian Map of the Soul in the Digital Age (Read Through Structural Intelligence)
by: Jovanovic, Vladisav
Published: (2026)
by: Jovanovic, Vladisav
Published: (2026)
The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation
by: Devin, Andrew James
Published: (2026)
by: Devin, Andrew James
Published: (2026)
I Let Claude Run My Fantasy Football Team for a Whole Season — It Beat 11 of My Friends
by: AI Angels
Published: (2026)
by: AI Angels
Published: (2026)
Why Voice-Mode Gemini Beat My $400 Italian Tutor in 21 Days (Full Daily Script Inside)
by: AI Angels
Published: (2026)
by: AI Angels
Published: (2026)
GALATEA II: Benchmarking LLM Safety in Clinical Simulation. Behavioural Safety and Ethical Robustness of Large Language Models in a Multi-Agent ICU Decision Support Architecture
by: Shlyakhta, Taras
Published: (2026)
by: Shlyakhta, Taras
Published: (2026)
Shallow Pass Budget Constraints and Structured Data Trade-offs in LLM Training Ingestion
by: Mas, Joseph
Published: (2026)
by: Mas, Joseph
Published: (2026)
AI Eval Forge: Mixed-Check Regression Testing for LLM and Agent Workflows
by: Katta, Mukunda Rao
Published: (2026)
by: Katta, Mukunda Rao
Published: (2026)
Executive Summary: AI Privacy Risks and Mitigations in Large Language Models
by: Khan, Masood
Published: (2025)
by: Khan, Masood
Published: (2025)
The Four-Layer Model: A Socio-Psychological Framework for LLM Behavior
by: Delannoy, Lorenzo, et al.
Published: (2026)
by: Delannoy, Lorenzo, et al.
Published: (2026)
aikenkyu001/iterative_self_healing_benchmark: v1.0.0: Scaffolding Trinity for Deterministic LLM Code Generation
by: Miyata, Fumio
Published: (2026)
by: Miyata, Fumio
Published: (2026)
Signs of Life - Visual Art from 469 Conversations with Claude
by: Chesterton, Bo, et al.
Published: (2026)
by: Chesterton, Bo, et al.
Published: (2026)
Machine-Readable Behavioural Compliance Evidence for AI Systems: A Specification Profiling Framework
by: Caprazli, Kafkas M.
Published: (2026)
by: Caprazli, Kafkas M.
Published: (2026)
Anti-Hydra vs Anthropic Benchmark Comparison
by: Ochej, Stephane
Published: (2026)
by: Ochej, Stephane
Published: (2026)
Oracle Difficulty Decomposed: Four Independent Mechanisms Explain 95%+ of Benchmark Variance
by: Sanchez, Bryan
Published: (2026)
by: Sanchez, Bryan
Published: (2026)
Deterministic σ-Regularized Benchmarking of the Cekirge Model Against GPT-Transformer Baselines
by: CEKIRGE, Huseyin Murat
Published: (2025)
by: CEKIRGE, Huseyin Murat
Published: (2025)
AI Writing Ethics: Responsible and Ethical Use of Generative AI in Academic Writing
by: Zuzafre, Mohd Nor, et al.
Published: (2026)
by: Zuzafre, Mohd Nor, et al.
Published: (2026)
OMNIA-MINIMAL: Structural Stability Beyond Surface Correctness
by: Brighindi, Massimiliano
Published: (2026)
by: Brighindi, Massimiliano
Published: (2026)
Pattern Pressure, Accuracy Drift, and False User-State Attribution
by: Honeycutt, Edwin Marshall III
Published: (2026)
by: Honeycutt, Edwin Marshall III
Published: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
by: Bernal Díaz, Víctor Cristóbal
Published: (2026)
by: Bernal Díaz, Víctor Cristóbal
Published: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001 V1.1. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
by: Bernal Díaz, Víctor Cristóbal
Published: (2026)
by: Bernal Díaz, Víctor Cristóbal
Published: (2026)
Supplementary materials for Words That Won't Hold Still
by: Reynolds, Brett
Published: (2025)
by: Reynolds, Brett
Published: (2025)
31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
by: Bernal Díaz, Víctor Cristóbal
Published: (2026)
by: Bernal Díaz, Víctor Cristóbal
Published: (2026)
Extra Large Language Models Benchmarking for Medicinal Chemistry
by: Kawchak, Kevin
Published: (2024)
by: Kawchak, Kevin
Published: (2024)
Selection Mechanics Framework (SMF): A Structural Perspective on Variability and Its Transformation in Large Language Model Outputs
by: Kaneda, Mutsumi
Published: (2026)
by: Kaneda, Mutsumi
Published: (2026)
Hidden in plain sight: Unraveling compensation disclosure bloat with generative AI and its impact on executive compensation
by: Burduli, Lizi, et al.
Published: (2026)
by: Burduli, Lizi, et al.
Published: (2026)
A Comparative Analysis of AI-Generated vs. Human-Written Marketing Content Using Readability and SEO Metrics
by: Abuhumaid, Ibrahim
Published: (2026)
by: Abuhumaid, Ibrahim
Published: (2026)
A Benchmark for Symbolic Reasoning from Pixel Sequences: Grid-Level Visual Completion and Correction
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
Language-as-Dimension Theory (LDT): A New Scientific Framework for Intelligence, Meaning, and Reality Formation
by: Woodard, Bethany, et al.
Published: (2025)
by: Woodard, Bethany, et al.
Published: (2025)
Similar Items
-
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
by: Khare, Mohit
Published: (2026) -
Gemini Update Clinical decision support based on Bevacizumab cancer trials and pushing the limitations of advanced LLMs
by: Kawchak, Kevin
Published: (2025) -
The Brain Problem: Creative Constraint Optimization in Large Language Models
by: Marinello, Nicola, et al.
Published: (2026) -
Estimating the Impact of Automation on Vocational Education: The Case of Technical Courses
by: Lima, Yuri, et al.
Published: (2024) -
LACF Emotional Paradigm: A Personalized Artificial Nervous System for Human-AI Alignment
by: Ochej, Stephane, et al.
Published: (2026)