Examples of metrics and diagnostics for different evaluation and benchmarking approaches
Fuente:
Zenodo
Saved in:
| Main Author: | Hassler, Birgit |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2025
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Different approaches that are most commonly used for the evaluation and benchmarking of climate models.
by: Lembo, Valerio, et al.
Published: (2024)
by: Lembo, Valerio, et al.
Published: (2024)
Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
A declarative approach and benchmark tool for controlled evaluation of microservice resiliency patterns
by: Carlos M. Aderaldo, et al.
Published: (2024)
by: Carlos M. Aderaldo, et al.
Published: (2024)
A practical generalization metric for deep networks benchmarking
by: Huang, Mengqing, et al.
Published: (2024)
by: Huang, Mengqing, et al.
Published: (2024)
Lower bounds on non-random fluctuations in planar first passage percolation
by: Hassler, Malte
Published: (2025)
by: Hassler, Malte
Published: (2025)
Classification of the limit shape for 1+1-dimensional FPP
by: Hassler, Malte
Published: (2024)
by: Hassler, Malte
Published: (2024)
Topological quantum computing
by: Hassler, Fabian
Published: (2024)
by: Hassler, Fabian
Published: (2024)
Ambivalenz der Wiedereingliederung
by: Hassler, Benedikt
Published: (2021)
by: Hassler, Benedikt
Published: (2021)
A bijection between edges of the Turán graph and irreducible elements in the dominance order lattice
by: Hassler, Nathanaël
Published: (2026)
by: Hassler, Nathanaël
Published: (2026)
Evidential and epistemic sentence adverbs in Romance languages
by: Gerda Haßler
Published: (2018)
by: Gerda Haßler
Published: (2018)
Changing Your Domain Name in 25 Nail-Biting Steps
by: Hassler, Carol
Published: (2012)
by: Hassler, Carol
Published: (2012)
A reproducible diagnostic benchmark for language-to-target generation in robotic manipulation
by: Chen, Zhuo
Published: (2026)
by: Chen, Zhuo
Published: (2026)
Examples of étale extensions of Green functors
by: Lindenstrauss, Ayelet, et al.
Published: (2023)
by: Lindenstrauss, Ayelet, et al.
Published: (2023)
A NATUREZA NA CIDADE: UMA ABORDAGEM A PARTIR DA PERCEPÇÃO DA POPULAÇÃO ACERCA DO JARDIM BOTÂNICO DE CURITIBA (PR)
by: Márcio Luís Hassler
Published: (2006)
by: Márcio Luís Hassler
Published: (2006)
CONTRIBUIÇÃO GEOGRÁFICA PARA O ESTUDO DO LUGAR
by: Márcio Luís Hassler
Published: (2009)
by: Márcio Luís Hassler
Published: (2009)
A IMPORTÂNCIA DAS UNIDADES DE CONSERVAÇÃO NO BRASIL
by: Márcio Luís Hassler
Published: (2005)
by: Márcio Luís Hassler
Published: (2005)
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
by: Ailem, Melissa, et al.
Published: (2024)
by: Ailem, Melissa, et al.
Published: (2024)
wuzezhen5577/Model-Fidelity-Metric: Model Fidelity Metric: A robust and diagnostic metric for land surface model evaluation
by: Wu Zezhen
Published: (2026)
by: Wu Zezhen
Published: (2026)
CSPBench: a benchmark and critical evaluation of Crystal Structure Prediction
by: Wei, Lai, et al.
Published: (2024)
by: Wei, Lai, et al.
Published: (2024)
How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation
by: Singh, Prabhant, et al.
Published: (2025)
by: Singh, Prabhant, et al.
Published: (2025)
Script collapse in multilingual ASR: A reference-free metric and 100-pair benchmark
by: Rahman, Hanif
Published: (2026)
by: Rahman, Hanif
Published: (2026)
All-order generalized Green-Schwarz transformations
by: Gitsis, Achilleas, et al.
Published: (2025)
by: Gitsis, Achilleas, et al.
Published: (2025)
Consistent truncation and generalized duality based on exceptional generalized cosets
by: Hassler, Falk, et al.
Published: (2025)
by: Hassler, Falk, et al.
Published: (2025)
Emerging consecutive pattern avoidance
by: Hassler, Nathanaël, et al.
Published: (2025)
by: Hassler, Nathanaël, et al.
Published: (2025)
Self‐Normalising Tests Using the Cauchy Distribution
by: Uwe Hassler, et al.
Published: (2025)
by: Uwe Hassler, et al.
Published: (2025)
Combinatorial Hubbard trees for postcritically infinite unicritical polynomials and exponential maps
by: Hassler, Malte, et al.
Published: (2024)
by: Hassler, Malte, et al.
Published: (2024)
Unraveling the generalized Bergshoeff-de Roo identification
by: Gitsis, Achilleas, et al.
Published: (2024)
by: Gitsis, Achilleas, et al.
Published: (2024)
Superbunched radiation of a tunnel junction due to charge quantization
by: Kim, Steven, et al.
Published: (2024)
by: Kim, Steven, et al.
Published: (2024)
Photon counting beyond the rotating-wave approximation
by: Kim, Steven, et al.
Published: (2026)
by: Kim, Steven, et al.
Published: (2026)
Hölder continuity of core entropy for non-recurrent quadratic polynomials
by: Hassler, Malte, et al.
Published: (2024)
by: Hassler, Malte, et al.
Published: (2024)
All maximal gauged supergravities with uplift
by: Hassler, Falk, et al.
Published: (2022)
by: Hassler, Falk, et al.
Published: (2022)
The American college of radiology diagnostic fluoroscopy dose index registry pilot: Dosimetric performance and benchmarking challenges
by: Steve D. Mann, et al.
Published: (2026)
by: Steve D. Mann, et al.
Published: (2026)
Resolving power: A general approach to compare the distinguishing ability of threshold-free evaluation metrics
by: Beam, Colin S.
Published: (2023)
by: Beam, Colin S.
Published: (2023)
Picoeukaryote counts determined from the community sampled at station West and East (PS97) before and after exposure to different Fe and Mn availabilities
by: Balaguer, Jenna, et al.
Published: (2021)
by: Balaguer, Jenna, et al.
Published: (2021)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
by: Montalan, Jann Railey, et al.
Published: (2025)
by: Montalan, Jann Railey, et al.
Published: (2025)
VWise: A novel benchmark for evaluating scene classification for vehicular applications
by: Azevedo, Pedro, et al.
Published: (2024)
by: Azevedo, Pedro, et al.
Published: (2024)
Phytoplankton species determination and counts determined from the community sampled at station West and East (PS97) before and after exposure to different Fe and Mn availabilities
by: Balaguer, Jenna, et al.
Published: (2021)
by: Balaguer, Jenna, et al.
Published: (2021)
Inconsistency of evaluation metrics in link prediction
by: Bi, Yilin, et al.
Published: (2024)
by: Bi, Yilin, et al.
Published: (2024)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
by: Li, Kunning, et al.
Published: (2025)
by: Li, Kunning, et al.
Published: (2025)
Extreme Weather Bench: A framework and benchmark for evaluation of high-impact weather
by: McGovern, Amy, et al.
Published: (2026)
by: McGovern, Amy, et al.
Published: (2026)
Similar Items
-
Different approaches that are most commonly used for the evaluation and benchmarking of climate models.
by: Lembo, Valerio, et al.
Published: (2024) -
Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
by: Chen, Hao, et al.
Published: (2025) -
A declarative approach and benchmark tool for controlled evaluation of microservice resiliency patterns
by: Carlos M. Aderaldo, et al.
Published: (2024) -
A practical generalization metric for deep networks benchmarking
by: Huang, Mengqing, et al.
Published: (2024) -
Lower bounds on non-random fluctuations in planar first passage percolation
by: Hassler, Malte
Published: (2025)