Branching Out: Broadening AI Measurement and Evaluation with Measurement Trees
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Greenberg, Craig, Hall, Patrick, Jensen, Theodore, Greene, Kristen, Amironesei, Razvan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open-World Evaluations for Measuring Frontier AI Capabilities
von: Kapoor, Sayash, et al.
Veröffentlicht: (2026)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2026)
Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems
von: Johnson, Rebecca L.
Veröffentlicht: (2026)
von: Johnson, Rebecca L.
Veröffentlicht: (2026)
Measuring the metacognition of AI
von: Servajean, Richard, et al.
Veröffentlicht: (2026)
von: Servajean, Richard, et al.
Veröffentlicht: (2026)
In between myth and reality: AI for math -- a case study in category theory
von: Diaconescu, Răzvan
Veröffentlicht: (2025)
von: Diaconescu, Răzvan
Veröffentlicht: (2025)
Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
von: Grey, Markov, et al.
Veröffentlicht: (2025)
von: Grey, Markov, et al.
Veröffentlicht: (2025)
StreetDesignAI: Broadening Designer Perspectives Through Multi-Persona Evaluation of Cycling Infrastructure
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
Measuring AI Alignment with Human Flourishing
von: Hilliard, Elizabeth, et al.
Veröffentlicht: (2025)
von: Hilliard, Elizabeth, et al.
Veröffentlicht: (2025)
Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation
von: von Däniken, Pius, et al.
Veröffentlicht: (2024)
von: von Däniken, Pius, et al.
Veröffentlicht: (2024)
Don't Measure Once: Measuring Visibility in AI Search (GEO)
von: Schulte, Julius, et al.
Veröffentlicht: (2026)
von: Schulte, Julius, et al.
Veröffentlicht: (2026)
Measuring What Matters: The AI Pluralism Index
von: Mushkani, Rashid
Veröffentlicht: (2025)
von: Mushkani, Rashid
Veröffentlicht: (2025)
Predicting Human Chess Moves: An AI Assisted Analysis of Chess Games Using Skill-group Specific n-gram Language Models
von: Zhong, Daren, et al.
Veröffentlicht: (2025)
von: Zhong, Daren, et al.
Veröffentlicht: (2025)
Measuring the environmental impact of delivering AI at Google Scale
von: Elsworth, Cooper, et al.
Veröffentlicht: (2025)
von: Elsworth, Cooper, et al.
Veröffentlicht: (2025)
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
von: Choong, Yee-Yin, et al.
Veröffentlicht: (2026)
von: Choong, Yee-Yin, et al.
Veröffentlicht: (2026)
Evaluating the Effects of AI Directors for Quest Selection
von: Yu, Kristen K., et al.
Veröffentlicht: (2024)
von: Yu, Kristen K., et al.
Veröffentlicht: (2024)
Measuring AI R&D Automation
von: Chan, Alan, et al.
Veröffentlicht: (2026)
von: Chan, Alan, et al.
Veröffentlicht: (2026)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
von: Testini, Irene, et al.
Veröffentlicht: (2025)
von: Testini, Irene, et al.
Veröffentlicht: (2025)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
von: Voudouris, Konstantinos, et al.
Veröffentlicht: (2026)
von: Voudouris, Konstantinos, et al.
Veröffentlicht: (2026)
Broadening Ontologization Design: Embracing Data Pipeline Strategies
von: Partridge, Chris, et al.
Veröffentlicht: (2025)
von: Partridge, Chris, et al.
Veröffentlicht: (2025)
Techniques for Measuring the Inferential Strength of Forgetting Policies
von: Doherty, Patrick, et al.
Veröffentlicht: (2024)
von: Doherty, Patrick, et al.
Veröffentlicht: (2024)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
von: Murugadoss, Bhuvanashree, et al.
Veröffentlicht: (2024)
von: Murugadoss, Bhuvanashree, et al.
Veröffentlicht: (2024)
Variance-Bounded Evaluation of Entity-Centric AI Systems Without Ground Truth: Theory and Measurement
von: Ding, Kaihua
Veröffentlicht: (2025)
von: Ding, Kaihua
Veröffentlicht: (2025)
Measuring AI Reasoning: A Guide for Researchers
von: Nwadike, Munachiso Samuel, et al.
Veröffentlicht: (2026)
von: Nwadike, Munachiso Samuel, et al.
Veröffentlicht: (2026)
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
von: Desikan, Prasanna, et al.
Veröffentlicht: (2026)
von: Desikan, Prasanna, et al.
Veröffentlicht: (2026)
InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
von: Wan, Yixin, et al.
Veröffentlicht: (2025)
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
von: Guo, Xuehang, et al.
Veröffentlicht: (2025)
von: Guo, Xuehang, et al.
Veröffentlicht: (2025)
Decomposing and Measuring Evaluation Awareness
von: Li, Changling, et al.
Veröffentlicht: (2026)
von: Li, Changling, et al.
Veröffentlicht: (2026)
Benchmark Transparency: Measuring the Impact of Data on Evaluation
von: Kovatchev, Venelin, et al.
Veröffentlicht: (2024)
von: Kovatchev, Venelin, et al.
Veröffentlicht: (2024)
Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources
von: Clark, Hannah-Beth, et al.
Veröffentlicht: (2025)
von: Clark, Hannah-Beth, et al.
Veröffentlicht: (2025)
LLM-Based Bot Broadens the Range of Arguments in Online Discussions, Even When Transparently Disclosed as AI
von: Vuk, Valeria, et al.
Veröffentlicht: (2025)
von: Vuk, Valeria, et al.
Veröffentlicht: (2025)
Measuring AI Ability to Complete Long Software Tasks
von: Kwa, Thomas, et al.
Veröffentlicht: (2025)
von: Kwa, Thomas, et al.
Veröffentlicht: (2025)
Feedback Forensics: A Toolkit to Measure AI Personality
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
von: Findeis, Arduin, et al.
Veröffentlicht: (2025)
LLM Rationalis? Measuring Bargaining Capabilities of AI Negotiators
von: Shah, Cheril, et al.
Veröffentlicht: (2025)
von: Shah, Cheril, et al.
Veröffentlicht: (2025)
Using AI to Measure Parkinson's Disease Severity at Home
von: Islam, Md Saiful, et al.
Veröffentlicht: (2023)
von: Islam, Md Saiful, et al.
Veröffentlicht: (2023)
Measuring AI agent autonomy: Towards a scalable approach with code inspection
von: Cihon, Peter, et al.
Veröffentlicht: (2025)
von: Cihon, Peter, et al.
Veröffentlicht: (2025)
Measuring the Measurers: Quality Evaluation of Hallucination Benchmarks for Large Vision-Language Models
von: Yan, Bei, et al.
Veröffentlicht: (2024)
von: Yan, Bei, et al.
Veröffentlicht: (2024)
Intentionality is a Design Decision: Measuring Functional Intentionality for Accountable AI Systems
von: Chiappetta, Allessia, et al.
Veröffentlicht: (2026)
von: Chiappetta, Allessia, et al.
Veröffentlicht: (2026)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks
von: Qu, Yubin, et al.
Veröffentlicht: (2026)
von: Qu, Yubin, et al.
Veröffentlicht: (2026)
The Evaluation Trap: Benchmark Design as Theoretical Commitment
von: Kalaitzidis, Theodore J
Veröffentlicht: (2026)
von: Kalaitzidis, Theodore J
Veröffentlicht: (2026)
DBMF: A Dual-Branch Multimodal Framework for Out-of-Distribution Detection
von: Yue, Jiangbei, et al.
Veröffentlicht: (2026)
von: Yue, Jiangbei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Open-World Evaluations for Measuring Frontier AI Capabilities
von: Kapoor, Sayash, et al.
Veröffentlicht: (2026) -
Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems
von: Johnson, Rebecca L.
Veröffentlicht: (2026) -
Measuring the metacognition of AI
von: Servajean, Richard, et al.
Veröffentlicht: (2026) -
In between myth and reality: AI for math -- a case study in category theory
von: Diaconescu, Răzvan
Veröffentlicht: (2025) -
Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
von: Grey, Markov, et al.
Veröffentlicht: (2025)