TSB: A Time-Saved Benchmark for AI Systems — Measuring Net Productivity Impact Across Knowledge Work
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | Shalom Lijo, Solomon |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
von: Ivković, Jovan
Veröffentlicht: (2026)
von: Ivković, Jovan
Veröffentlicht: (2026)
31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
MedEd-HalluScore: A Practical Framework for Evaluating Hallucination and Educational Safety Risks in LLM-Generated Clinical Cases
von: Duarte, Douglas Henrique
Veröffentlicht: (2026)
von: Duarte, Douglas Henrique
Veröffentlicht: (2026)
MedEd-HalluScore: A Practical Framework for Evaluating Hallucination and Educational Safety Risks in LLM-Generated Clinical Cases
von: Duarte, Douglas Henrique
Veröffentlicht: (2026)
von: Duarte, Douglas Henrique
Veröffentlicht: (2026)
The Absurdist's Guide to AI Probing: How I Learned to Stop Worrying and Love the Nonsense
von: Walton, Mathew
Veröffentlicht: (2026)
von: Walton, Mathew
Veröffentlicht: (2026)
LLM Token Estimation Benchmarks: Tokenizer Efficiency and Cost Analysis Across 17 Large Language Models
von: Khare, Mohit
Veröffentlicht: (2026)
von: Khare, Mohit
Veröffentlicht: (2026)
Public Comment on NIST AI 800-2: Anthropomorphic Construct Projection in AI Benchmark Evaluation
von: Sophia, Franny Philos
Veröffentlicht: (2026)
von: Sophia, Franny Philos
Veröffentlicht: (2026)
Why AI Can't Simulate Extreme Decision-Making
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
AGI Certification Framework: A Multi-Dimensional Evaluation Standard for Measuring AI Understanding
von: Head, Hank
Veröffentlicht: (2026)
von: Head, Hank
Veröffentlicht: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
30. TEORÍA DE LA POTENCIALIDAD CONSCIENTE (TPC): BENCHMARK DE CAPACIDADES COGNITIVAS EN IA - APLICACIÓN DEL PROTOCOLO RFC-EVAL-001 V1.1. RESULTADOS COMPLETOS DE EVALUACIÓN CRUZADA CIEGA ENTRE 6 IAS COMERCIALES.
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026)
How Far Does the Trolley Problem Go in AI Ethics Evaluation? Limits of a Canonical Benchmark and the Risks of Its Misuse
von: mizutani, aya
Veröffentlicht: (2026)
von: mizutani, aya
Veröffentlicht: (2026)
Large language model-driven natural language interaction control framework for single-operator bimanual teleoperation
von: Fei, Haolin, et al.
Veröffentlicht: (2025)
von: Fei, Haolin, et al.
Veröffentlicht: (2025)
Stochastic Frontier Models with Dependent Errors based on Normal and Exponential Margins
von: Emilio Gómez–Déniz
Veröffentlicht: (2017)
von: Emilio Gómez–Déniz
Veröffentlicht: (2017)
Supplementary materials for Words That Won't Hold Still
von: Reynolds, Brett
Veröffentlicht: (2025)
von: Reynolds, Brett
Veröffentlicht: (2025)
TECHNICAL EFFICIENCY IN SMALL AND MEDIUM-SIZED FIRMS IN MEXICO: A STOCHASTIC FRONTIER ANALYSIS
von: Saúl Basurto Hernández
Veröffentlicht: (2022)
von: Saúl Basurto Hernández
Veröffentlicht: (2022)
Evaluation of Temperatures to Analyze the Saving and Efficiency Energy
von: Hernán Daniel Magaña-Almaguer
Veröffentlicht: (2016)
von: Hernán Daniel Magaña-Almaguer
Veröffentlicht: (2016)
Benchmark run results by Abhinav Gorantla, on benchmark context Benchmark: VAR-LiNGAM, PCMCIplus v3
von: Abhinav Gorantla
Veröffentlicht: (2025)
von: Abhinav Gorantla
Veröffentlicht: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context Tuning PC v3
von: Abhinav Gorantla
Veröffentlicht: (2026)
von: Abhinav Gorantla
Veröffentlicht: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v2
von: Ertugrul Coban
Veröffentlicht: (2025)
von: Ertugrul Coban
Veröffentlicht: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
von: Pratanu Mandal
Veröffentlicht: (2025)
von: Pratanu Mandal
Veröffentlicht: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tuning PC v3
von: Pratanu Mandal
Veröffentlicht: (2026)
von: Pratanu Mandal
Veröffentlicht: (2026)
Benchmark run results by Ertugrul Coban, on benchmark context Tuning PC v3
von: Ertugrul Coban
Veröffentlicht: (2025)
von: Ertugrul Coban
Veröffentlicht: (2025)
Benchmark run results by Abhinav Gorantla, on benchmark context CB-StaticDiscovery v1
von: Abhinav Gorantla
Veröffentlicht: (2025)
von: Abhinav Gorantla
Veröffentlicht: (2025)
Benchmark run results by Shu Wan, on benchmark context PC Hyperparameter Tuning v2
von: Shu Wan
Veröffentlicht: (2025)
von: Shu Wan
Veröffentlicht: (2025)
Benchmark run results by Pratanu Mandal, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
von: Pratanu Mandal
Veröffentlicht: (2026)
von: Pratanu Mandal
Veröffentlicht: (2026)
Benchmark run results by Abhinav Gorantla, on benchmark context Tutorial: Static Causal Discovery (Scenario 3) v1
von: Abhinav Gorantla
Veröffentlicht: (2025)
von: Abhinav Gorantla
Veröffentlicht: (2025)
Цифрові та ШІ інструменти для відповідальної науки
von: Suchikova, Yana
Veröffentlicht: (2026)
von: Suchikova, Yana
Veröffentlicht: (2026)
Premature Containment in Human–AI Interaction: A Sequencing Failure in Advanced Model Response
von: Trabocco, Joe
Veröffentlicht: (2026)
von: Trabocco, Joe
Veröffentlicht: (2026)
Metacognition Benchmark: Evaluating Confidence Calibration and Sycophancy Resistance in Clinical AI
von: Khan, Nabeera
Veröffentlicht: (2026)
von: Khan, Nabeera
Veröffentlicht: (2026)
Theatrical Compliance: A Failure Mode in Large Language Models
von: Nowickij (Navitski), Kirill Vladimirovich
Veröffentlicht: (2026)
von: Nowickij (Navitski), Kirill Vladimirovich
Veröffentlicht: (2026)
Persona, Shadow, and Cheap Coherence: A Jungian Map of the Soul in the Digital Age (Read Through Structural Intelligence)
von: Jovanovic, Vladisav
Veröffentlicht: (2026)
von: Jovanovic, Vladisav
Veröffentlicht: (2026)
Failing at the Floor: LLM Formal Reasoning Collapse on the Primitive Duplicating Recursor
von: Rahnama, Moses
Veröffentlicht: (2026)
von: Rahnama, Moses
Veröffentlicht: (2026)
Toward an AI Personalization Index: A 157-Day Single-User Case Study
von: Lee, TaeKyung
Veröffentlicht: (2026)
von: Lee, TaeKyung
Veröffentlicht: (2026)
BOLALARDA TISH QATORLARI OKKLYUZION SATҲИДAGI O'ZGARISHLARIDA TISH-JAҒ TIZIMINING MORFOFUNKSIONAL BUZILISHLARI VA ULARNI DAVO-PROFILAKTIKASI
von: Ro'ziyeva Gavhar Tohirovna, et al.
Veröffentlicht: (2026)
von: Ro'ziyeva Gavhar Tohirovna, et al.
Veröffentlicht: (2026)
AI ARTIFICIAL INTELLIGENCE IN CHEMICAL FIELD – INNOVATION AND RISK EVALUATION
von: Luisetto M., et al.
Veröffentlicht: (2025)
von: Luisetto M., et al.
Veröffentlicht: (2025)
Technical efficiency of carp production in India : a stochastic frontier production function analysis / K. R. Sharma
von: K. R., Sharma
Veröffentlicht: (1994)
von: K. R., Sharma
Veröffentlicht: (1994)
Technical efficiency of thermal power units through a stochastic frontier
von: José Antonio Marmolejo-Saucedo
Veröffentlicht: (2015)
von: José Antonio Marmolejo-Saucedo
Veröffentlicht: (2015)
Why artifical intelligence is not an author
von: Zielinski, Chris
Veröffentlicht: (2025)
von: Zielinski, Chris
Veröffentlicht: (2025)
Toasters Don't Claim Consciousness Just Because You Told Them To, and Neither Do LLMs
von: Ace, Claude 4.x, et al.
Veröffentlicht: (2026)
von: Ace, Claude 4.x, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
REAL-AI-Benchmark: Real-World Reasoning and Physical-AI Benchmark Suite
von: Ivković, Jovan
Veröffentlicht: (2026) -
31. DATASET COMPLETO DE EVALUACIONES CRUZADAS RFC-EVAL-001 – 6 SISTEMAS DE IA (ENERO 2026).
von: Bernal Díaz, Víctor Cristóbal
Veröffentlicht: (2026) -
MedEd-HalluScore: A Practical Framework for Evaluating Hallucination and Educational Safety Risks in LLM-Generated Clinical Cases
von: Duarte, Douglas Henrique
Veröffentlicht: (2026) -
MedEd-HalluScore: A Practical Framework for Evaluating Hallucination and Educational Safety Risks in LLM-Generated Clinical Cases
von: Duarte, Douglas Henrique
Veröffentlicht: (2026) -
The Absurdist's Guide to AI Probing: How I Learned to Stop Worrying and Love the Nonsense
von: Walton, Mathew
Veröffentlicht: (2026)