Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
Fuente:
Zenodo
Enregistré dans:
| Auteur principal: | Khan, Nabeera |
|---|---|
| Format: | Recurso digital |
| Langue: | anglais |
| Publié: |
Zenodo
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
par: Khare, Mohit
Publié: (2026)
par: Khare, Mohit
Publié: (2026)
Gemini Update Clinical decision support based on Bevacizumab cancer trials and pushing the limitations of advanced LLMs
par: Kawchak, Kevin
Publié: (2025)
par: Kawchak, Kevin
Publié: (2025)
The Brain Problem: Creative Constraint Optimization in Large Language Models
par: Marinello, Nicola, et autres
Publié: (2026)
par: Marinello, Nicola, et autres
Publié: (2026)
Estimating the Impact of Automation on Vocational Education: The Case of Technical Courses
par: Lima, Yuri, et autres
Publié: (2024)
par: Lima, Yuri, et autres
Publié: (2024)
LACF Emotional Paradigm: A Personalized Artificial Nervous System for Human-AI Alignment
par: Ochej, Stephane, et autres
Publié: (2026)
par: Ochej, Stephane, et autres
Publié: (2026)
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
par: Head, Hank
Publié: (2026)
par: Head, Hank
Publié: (2026)
Public Comment on NIST AI 800-2: Anthropomorphic Construct Projection in AI Benchmark Evaluation
par: Sophia, Franny Philos
Publié: (2026)
par: Sophia, Franny Philos
Publié: (2026)
When AI Tells You What You Want to Hear: Sycophantic Behavior of Large Language Models in Dementia Care Settings
par: Kolb, Christian
Publié: (2026)
par: Kolb, Christian
Publié: (2026)
Anima AI Community Pulse Dataset
par: AI Companion Picker, et autres
Publié: (2026)
par: AI Companion Picker, et autres
Publié: (2026)
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
par: Ivković, Jovan
Publié: (2026)
par: Ivković, Jovan
Publié: (2026)
Theatrical Compliance: A Failure Mode in Large Language Models
par: Nowickij (Navitski), Kirill Vladimirovich
Publié: (2026)
par: Nowickij (Navitski), Kirill Vladimirovich
Publié: (2026)
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
par: mizutani, aya
Publié: (2026)
par: mizutani, aya
Publié: (2026)
Persona, Shadow, and Cheap Coherence: A Jungian Map of the Soul in the Digital Age (Read Through Structural Intelligence)
par: Jovanovic, Vladisav
Publié: (2026)
par: Jovanovic, Vladisav
Publié: (2026)
The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation
par: Devin, Andrew James
Publié: (2026)
par: Devin, Andrew James
Publié: (2026)
I Let Claude Run My Fantasy Football Team for a Whole Season — It Beat 11 of My Friends
par: AI Angels
Publié: (2026)
par: AI Angels
Publié: (2026)
Why Voice-Mode Gemini Beat My $400 Italian Tutor in 21 Days (Full Daily Script Inside)
par: AI Angels
Publié: (2026)
par: AI Angels
Publié: (2026)
GALATEA II: Benchmarking LLM Safety in Clinical Simulation. Behavioural Safety and Ethical Robustness of Large Language Models in a Multi-Agent ICU Decision Support Architecture
par: Shlyakhta, Taras
Publié: (2026)
par: Shlyakhta, Taras
Publié: (2026)
Shallow Pass Budget Constraints and Structured Data Trade-offs in LLM Training Ingestion
par: Mas, Joseph
Publié: (2026)
par: Mas, Joseph
Publié: (2026)
AI Eval Forge: Mixed-Check Regression Testing for LLM and Agent Workflows
par: Katta, Mukunda Rao
Publié: (2026)
par: Katta, Mukunda Rao
Publié: (2026)
Executive Summary: AI Privacy Risks and Mitigations in Large Language Models
par: Khan, Masood
Publié: (2025)
par: Khan, Masood
Publié: (2025)
The Four-Layer Model: A Socio-Psychological Framework for LLM Behavior
par: Delannoy, Lorenzo, et autres
Publié: (2026)
par: Delannoy, Lorenzo, et autres
Publié: (2026)
aikenkyu001/iterative_self_healing_benchmark: v1.0.0: Scaffolding Trinity for Deterministic LLM Code Generation
par: Miyata, Fumio
Publié: (2026)
par: Miyata, Fumio
Publié: (2026)
Signs of Life - Visual Art from 469 Conversations with Claude
par: Chesterton, Bo, et autres
Publié: (2026)
par: Chesterton, Bo, et autres
Publié: (2026)
Machine-Readable Behavioural Compliance Evidence for AI Systems: A Specification Profiling Framework
par: Caprazli, Kafkas M.
Publié: (2026)
par: Caprazli, Kafkas M.
Publié: (2026)
Anti-Hydra vs Anthropic Benchmark Comparison
par: Ochej, Stephane
Publié: (2026)
par: Ochej, Stephane
Publié: (2026)
Oracle Difficulty Decomposed: Four Independent Mechanisms Explain 95%+ of Benchmark Variance
par: Sanchez, Bryan
Publié: (2026)
par: Sanchez, Bryan
Publié: (2026)
Deterministic σ-Regularized Benchmarking of the Cekirge Model Against GPT-Transformer Baselines
par: CEKIRGE, Huseyin Murat
Publié: (2025)
par: CEKIRGE, Huseyin Murat
Publié: (2025)
AI Writing Ethics: Responsible and Ethical Use of Generative AI in Academic Writing
par: Zuzafre, Mohd Nor, et autres
Publié: (2026)
par: Zuzafre, Mohd Nor, et autres
Publié: (2026)
OMNIA-MINIMAL: Structural Stability Beyond Surface Correctness
par: Brighindi, Massimiliano
Publié: (2026)
par: Brighindi, Massimiliano
Publié: (2026)
Pattern Pressure, Accuracy Drift, and False User-State Attribution
par: Honeycutt, Edwin Marshall III
Publié: (2026)
par: Honeycutt, Edwin Marshall III
Publié: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001 V1.1. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
Supplementary materials for Words That Won't Hold Still
par: Reynolds, Brett
Publié: (2025)
par: Reynolds, Brett
Publié: (2025)
31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
Extra Large Language Models Benchmarking for Medicinal Chemistry
par: Kawchak, Kevin
Publié: (2024)
par: Kawchak, Kevin
Publié: (2024)
Selection Mechanics Framework (SMF): A Structural Perspective on Variability and Its Transformation in Large Language Model Outputs
par: Kaneda, Mutsumi
Publié: (2026)
par: Kaneda, Mutsumi
Publié: (2026)
Hidden in plain sight: Unraveling compensation disclosure bloat with generative AI and its impact on executive compensation
par: Burduli, Lizi, et autres
Publié: (2026)
par: Burduli, Lizi, et autres
Publié: (2026)
A Comparative Analysis of AI-Generated vs. Human-Written Marketing Content Using Readability and SEO Metrics
par: Abuhumaid, Ibrahim
Publié: (2026)
par: Abuhumaid, Ibrahim
Publié: (2026)
A Benchmark for Symbolic Reasoning from Pixel Sequences: Grid-Level Visual Completion and Correction
par: Kang, Lei, et autres
Publié: (2025)
par: Kang, Lei, et autres
Publié: (2025)
Language-as-Dimension Theory (LDT): A New Scientific Framework for Intelligence, Meaning, and Reality Formation
par: Woodard, Bethany, et autres
Publié: (2025)
par: Woodard, Bethany, et autres
Publié: (2025)
Documents similaires
-
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
par: Khare, Mohit
Publié: (2026) -
Gemini Update Clinical decision support based on Bevacizumab cancer trials and pushing the limitations of advanced LLMs
par: Kawchak, Kevin
Publié: (2025) -
The Brain Problem: Creative Constraint Optimization in Large Language Models
par: Marinello, Nicola, et autres
Publié: (2026) -
Estimating the Impact of Automation on Vocational Education: The Case of Technical Courses
par: Lima, Yuri, et autres
Publié: (2024) -
LACF Emotional Paradigm: A Personalized Artificial Nervous System for Human-AI Alignment
par: Ochej, Stephane, et autres
Publié: (2026)