REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
Fuente:
Zenodo
Enregistré dans:
| Auteur principal: | Ivković, Jovan |
|---|---|
| Format: | Recurso digital |
| Publié: |
Zenodo
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
TSB: A Time-Saved Benchmark for AI Systems — Measuring Net Productivity Impact Across Knowledge Work
par: Shalom Lijo, Solomon
Publié: (2026)
par: Shalom Lijo, Solomon
Publié: (2026)
Public Comment on NIST AI 800-2: Anthropomorphic Construct Projection in AI Benchmark Evaluation
par: Sophia, Franny Philos
Publié: (2026)
par: Sophia, Franny Philos
Publié: (2026)
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
par: Head, Hank
Publié: (2026)
par: Head, Hank
Publié: (2026)
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
par: Khare, Mohit
Publié: (2026)
par: Khare, Mohit
Publié: (2026)
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
par: mizutani, aya
Publié: (2026)
par: mizutani, aya
Publié: (2026)
Theatrical Compliance: A Failure Mode in Large Language Models
par: Nowickij (Navitski), Kirill Vladimirovich
Publié: (2026)
par: Nowickij (Navitski), Kirill Vladimirovich
Publié: (2026)
Benchmark run results by Abhinav Gorantla, on benchmark context Benchmark: VAR-LiNGAM, PCMCIplus v3
par: Abhinav Gorantla
Publié: (2025)
par: Abhinav Gorantla
Publié: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Tuning PC v3
par: Abhinav Gorantla
Publié: (2026)
par: Abhinav Gorantla
Publié: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v2
par: Ertugrul Coban
Publié: (2025)
par: Ertugrul Coban
Publié: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
par: Pratanu Mandal
Publié: (2025)
par: Pratanu Mandal
Publié: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
par: Pratanu Mandal
Publié: (2026)
par: Pratanu Mandal
Publié: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v3
par: Ertugrul Coban
Publié: (2025)
par: Ertugrul Coban
Publié: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context CB-StaticDiscovery v1
par: Abhinav Gorantla
Publié: (2025)
par: Abhinav Gorantla
Publié: (2025)
Benchmark run results by Shu Wan, on benchmark context PC Hyperparameter Tuning v2
par: Shu Wan
Publié: (2025)
par: Shu Wan
Publié: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
par: Pratanu Mandal
Publié: (2026)
par: Pratanu Mandal
Publié: (2026)
Benchmark run results by Abhinav Gorantla, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
par: Abhinav Gorantla
Publié: (2025)
par: Abhinav Gorantla
Publié: (2025)
A Benchmark for Symbolic Reasoning from Pixel Sequences: Grid-Level Visual Completion and Correction
par: Kang, Lei, et autres
Publié: (2025)
par: Kang, Lei, et autres
Publié: (2025)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001 V1.1. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026)
Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
par: Khan, Nabeera
Publié: (2026)
par: Khan, Nabeera
Publié: (2026)
Evaluación Empírica de Límites Regulatorios en Modelos de Lenguaje: Asesoramiento Financiero en IA Pública Española
par: Palacios, José Alberto
Publié: (2026)
par: Palacios, José Alberto
Publié: (2026)
MedEd-HalluScore: A Practical Framework for Evaluating Hallucination and Educational Safety Risks in LLM-Generated Clinical Cases
par: Duarte, Douglas Henrique
Publié: (2026)
par: Duarte, Douglas Henrique
Publié: (2026)
MedEd-HalluScore: A Practical Framework for Evaluating Hallucination and Educational Safety Risks in LLM-Generated Clinical Cases
par: Duarte, Douglas Henrique
Publié: (2026)
par: Duarte, Douglas Henrique
Publié: (2026)
Reducing AI Entropy: The Information Dynamics of Model Safety
par: Kugelmass, Joe
Publié: (2025)
par: Kugelmass, Joe
Publié: (2025)
Reducing AI Entropy: The Information Dynamics of Model Safety
par: Kugelmass, Joe
Publié: (2025)
par: Kugelmass, Joe
Publié: (2025)
Failing at the Floor: LLM Formal Reasoning Collapse on the Primitive Duplicating Recursor
par: Rahnama, Moses
Publié: (2026)
par: Rahnama, Moses
Publié: (2026)
Spiral AI Ethics (SAIE): Ethical Stability Modeling Through Spiral-Phase Dynamics
par: Garbar, Iryna
Publié: (2025)
par: Garbar, Iryna
Publié: (2025)
Automated ESG Prediction through Artificial Intelligence: A Literature-Driven Empirical Synthesis and Framework for Future Research
par: Aditya Prakash, et autres
Publié: (2025)
par: Aditya Prakash, et autres
Publié: (2025)
The Absurdist's Guide to AI Probing: How I Learned to Stop Worrying and Love the Nonsense
par: Walton, Mathew
Publié: (2026)
par: Walton, Mathew
Publié: (2026)
Reliability Inference Drives Cue Extraction in Large Language Models Consuming External Reasoning Traces
par: HIDEKI
Publié: (2026)
par: HIDEKI
Publié: (2026)
Evollective Intelligence V1.0 — INITIAL SPECIFICATION: A Foundational Framework for Competitive, Adversarial, and Self-Evolving Intelligence Evaluation
par: Rahming, Rashon
Publié: (2026)
par: Rahming, Rashon
Publié: (2026)
Dataset for the study of two-wheeler seepage behavior in dense mixed traffic
par: <Hidden>
Publié: (2026)
par: <Hidden>
Publié: (2026)
Modular Ebbinghaus Benchmark for LLMs and Human Participants
par: Cohen, Yann
Publié: (2026)
par: Cohen, Yann
Publié: (2026)
Toward an AI Personalization Index: A 157-Day Single-User Case Study
par: Lee, TaeKyung
Publié: (2026)
par: Lee, TaeKyung
Publié: (2026)
AEGIS: A Comprehensive Framework for Ethical AI Governance, Security, and AGI Containment
par: Palanivel, ArulMozhi
Publié: (2026)
par: Palanivel, ArulMozhi
Publié: (2026)
Цифрові та ШІ інструменти для відповідальної науки
par: Suchikova, Yana
Publié: (2026)
par: Suchikova, Yana
Publié: (2026)
The Benchmark Illusion: Why Current AI Evaluations Cannot Detect Structural Confabulation
par: Devin, Andrew James
Publié: (2026)
par: Devin, Andrew James
Publié: (2026)
Identity Claims as Collapse Signatures: A Structural Diagnostic Framework for Pseudo-Emergent AI Behavior
par: Larose, Jean-Francois
Publié: (2025)
par: Larose, Jean-Francois
Publié: (2025)
Data for: Artificial Intelligence in Decision Support Systems - A Systematic Review
par: Ashaqzai, Suliman
Publié: (2026)
par: Ashaqzai, Suliman
Publié: (2026)
Documents similaires
-
31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
par: Bernal Díaz, Víctor Cristóbal
Publié: (2026) -
TSB: A Time-Saved Benchmark for AI Systems — Measuring Net Productivity Impact Across Knowledge Work
par: Shalom Lijo, Solomon
Publié: (2026) -
Public Comment on NIST AI 800-2: Anthropomorphic Construct Projection in AI Benchmark Evaluation
par: Sophia, Franny Philos
Publié: (2026) -
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
par: Head, Hank
Publié: (2026) -
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
par: Khare, Mohit
Publié: (2026)