Unsteady Metrics and Benchmarking Cultures of AI Model Builders
Fuente:
arXiv
Saved in:
| Main Authors: | Baack, Stefan, Buschek, Christo, Bohacek, Maty |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Impact of AI-Generated Text on the Internet
by: Dolezal, Jonas, et al.
Published: (2026)
by: Dolezal, Jonas, et al.
Published: (2026)
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
by: Bohacek, Maty, et al.
Published: (2025)
by: Bohacek, Maty, et al.
Published: (2025)
Nepotistically Trained Generative-AI Models Collapse
by: Bohacek, Matyas, et al.
Published: (2023)
by: Bohacek, Matyas, et al.
Published: (2023)
Human Action CLIPs: Detecting AI-generated Human Motion
by: Bohacek, Matyas, et al.
Published: (2024)
by: Bohacek, Matyas, et al.
Published: (2024)
Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
GenAI Confessions: Black-box Membership Inference for Generative Image Models
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Can Pose Transfer Models Generate Realistic Human Motion?
by: Knapp, Vaclav, et al.
Published: (2025)
by: Knapp, Vaclav, et al.
Published: (2025)
The DeepSpeak Dataset
by: Barrington, Sarah, et al.
Published: (2024)
by: Barrington, Sarah, et al.
Published: (2024)
Synthetic Human Action Video Data Generation with Pose Transfer
by: Knapp, Vaclav, et al.
Published: (2025)
by: Knapp, Vaclav, et al.
Published: (2025)
Towards an AI-Driven Video-Based American Sign Language Dictionary: Exploring Design and Usage Experience with Learners
by: Hassan, Saad, et al.
Published: (2025)
by: Hassan, Saad, et al.
Published: (2025)
Assisted Debate Builder with Large Language Models
by: Faugier, Elliot, et al.
Published: (2024)
by: Faugier, Elliot, et al.
Published: (2024)
AutoBaxBuilder: Bootstrapping Code Security Benchmarking
by: von Arx, Tobias, et al.
Published: (2025)
by: von Arx, Tobias, et al.
Published: (2025)
Exploring the Lands Between: A Method for Finding Differences between AI-Decisions and Human Ratings through Generated Samples
by: Mecke, Lukas, et al.
Published: (2024)
by: Mecke, Lukas, et al.
Published: (2024)
Dataset of News Articles with Provenance Metadata for Media Relevance Assessment
by: Peterka, Tomas, et al.
Published: (2025)
by: Peterka, Tomas, et al.
Published: (2025)
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
by: Li, Tianle, et al.
Published: (2024)
by: Li, Tianle, et al.
Published: (2024)
BuilderBench: The Building Blocks of Intelligent Agents
by: Ghugare, Raj, et al.
Published: (2025)
by: Ghugare, Raj, et al.
Published: (2025)
ABodyBuilder3: Improved and scalable antibody structure predictions
by: Kenlay, Henry, et al.
Published: (2024)
by: Kenlay, Henry, et al.
Published: (2024)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Explaining Too Much? Understanding How Large Language Model Reasoning Traces Influence Performance and Metacognition
by: Fernandes, Daniela, et al.
Published: (2026)
by: Fernandes, Daniela, et al.
Published: (2026)
Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework
by: Zietsman, Christo
Published: (2026)
by: Zietsman, Christo
Published: (2026)
FMEA Builder: Expert Guided Text Generation for Equipment Maintenance
by: Lynch, Karol, et al.
Published: (2024)
by: Lynch, Karol, et al.
Published: (2024)
AISafetyBenchExplorer: A Metric-Aware Catalogue of AI Safety Benchmarks Reveals Fragmented Measurement and Weak Benchmark Governance
by: Solanke, Abiodun A.
Published: (2026)
by: Solanke, Abiodun A.
Published: (2026)
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
Physics Context Builders: A Modular Framework for Physical Reasoning in Vision-Language Models
by: Balazadeh, Vahid, et al.
Published: (2024)
by: Balazadeh, Vahid, et al.
Published: (2024)
"Personal Portfolio Builder Using MERN Stack With AI Integration"
by: Kalva, Ajay Kumar, et al.
Published: (2025)
by: Kalva, Ajay Kumar, et al.
Published: (2025)
DockSmith: Scaling Reliable Coding Environments via an Agentic Docker Builder
by: Zhang, Jiaran, et al.
Published: (2026)
by: Zhang, Jiaran, et al.
Published: (2026)
AgentBuilder: Exploring Scaffolds for Prototyping User Experiences of Interface Agents
by: Liang, Jenny T., et al.
Published: (2025)
by: Liang, Jenny T., et al.
Published: (2025)
AI-CARE: Carbon-Aware Reporting Evaluation Metric for AI Models
by: Santosh, KC, et al.
Published: (2026)
by: Santosh, KC, et al.
Published: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation
by: Metcalf, Sara, et al.
Published: (2026)
by: Metcalf, Sara, et al.
Published: (2026)
Conceptual Cultural Index: A Metric for Cultural Specificity via Relative Generality
by: Ohashi, Takumi, et al.
Published: (2026)
by: Ohashi, Takumi, et al.
Published: (2026)
From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making
by: Lee, Min Hun
Published: (2026)
by: Lee, Min Hun
Published: (2026)
Apprentice Tutor Builder: A Platform For Users to Create and Personalize Intelligent Tutors
by: Smith, Glen, et al.
Published: (2024)
by: Smith, Glen, et al.
Published: (2024)
LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations
by: Pham, Viet-Thanh, et al.
Published: (2026)
by: Pham, Viet-Thanh, et al.
Published: (2026)
An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics
by: Liu, Miri, et al.
Published: (2026)
by: Liu, Miri, et al.
Published: (2026)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
by: Demchak, Nathaniel, et al.
Published: (2024)
by: Demchak, Nathaniel, et al.
Published: (2024)
Safer Builders, Risky Maintainers: A Comparative Study of Breaking Changes in Human vs Agentic PRs
by: Ferdous, K M, et al.
Published: (2026)
by: Ferdous, K M, et al.
Published: (2026)
Positive Alignment: Artificial Intelligence for Human Flourishing
by: Laukkonen, Ruben, et al.
Published: (2026)
by: Laukkonen, Ruben, et al.
Published: (2026)
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia
by: Ayash, Lama, et al.
Published: (2025)
by: Ayash, Lama, et al.
Published: (2025)
Similar Items
-
The Impact of AI-Generated Text on the Internet
by: Dolezal, Jonas, et al.
Published: (2026) -
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
by: Bohacek, Maty, et al.
Published: (2025) -
Nepotistically Trained Generative-AI Models Collapse
by: Bohacek, Matyas, et al.
Published: (2023) -
Human Action CLIPs: Detecting AI-generated Human Motion
by: Bohacek, Matyas, et al.
Published: (2024) -
Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
by: Bohacek, Matyas, et al.
Published: (2025)