AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
Fuente:
arXiv
Saved in:
| Main Authors: | Hardy, Michael, Reuel, Anka, Zhang, Lijin, Casabianca, Jodi M., Truong, Sang, Dave, Yash, Lee, Hansol, Domingue, Benjamin, Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Bayesian Item Response Model with Decomposition (D-BIRD): Modeling Cohort and Individual Learning Over Time
by: Lee, Hansol, et al.
Published: (2025)
by: Lee, Hansol, et al.
Published: (2025)
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Generative AI Needs Adaptive Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Audit Cards: Contextualizing AI Evaluations
by: Staufer, Leon, et al.
Published: (2025)
by: Staufer, Leon, et al.
Published: (2025)
A Detailed Historical and Statistical Analysis of the Influence of Hardware Artifacts on SPEC Integer Benchmark Performance
by: Wang, Yueyao, et al.
Published: (2024)
by: Wang, Yueyao, et al.
Published: (2024)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
by: Haupt, Andreas, et al.
Published: (2026)
by: Haupt, Andreas, et al.
Published: (2026)
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Demystifying AI in Criminal Justice
by: Berk, Richard
Published: (2025)
by: Berk, Richard
Published: (2025)
RestoreAI -- Pattern-based Risk Estimation Of Remaining Explosives
by: Kischelewski, Björn, et al.
Published: (2025)
by: Kischelewski, Björn, et al.
Published: (2025)
Ball path curvature and in-game free throw shooting proficiency in the National Basketball Association
by: Zhu, Ruoqian, et al.
Published: (2025)
by: Zhu, Ruoqian, et al.
Published: (2025)
High-Dimensional Markov-switching Ordinary Differential Processes
by: Tsai, Katherine, et al.
Published: (2024)
by: Tsai, Katherine, et al.
Published: (2024)
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
by: Hardy, Amelia, et al.
Published: (2024)
by: Hardy, Amelia, et al.
Published: (2024)
Fairness in Reinforcement Learning: A Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Quantifying the AI Gap: A Comparative Index of Development in the United States and Chinese Regions
by: Li, Yuanxi, et al.
Published: (2025)
by: Li, Yuanxi, et al.
Published: (2025)
Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
by: Ramos, Mark Louie F.
Published: (2026)
by: Ramos, Mark Louie F.
Published: (2026)
Generative AI as a Safety Net for Survey Question Refinement
by: Metheney, Erica Ann, et al.
Published: (2025)
by: Metheney, Erica Ann, et al.
Published: (2025)
How do transportation professionals perceive the impacts of AI applications in transportation? A latent class cluster analysis
by: Qian, Yiheng, et al.
Published: (2024)
by: Qian, Yiheng, et al.
Published: (2024)
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
Exploring Student Interactions with AI-Powered Learning Tools: A Qualitative Study Connecting Interaction Patterns to Educational Learning Theories
by: Muzumdar, Prathamesh, et al.
Published: (2025)
by: Muzumdar, Prathamesh, et al.
Published: (2025)
Mapping Socio-Economic Divides with Urban Mobility Data
by: Liu, Yingche, et al.
Published: (2025)
by: Liu, Yingche, et al.
Published: (2025)
The Hidden AI Race: Tracking Environmental Costs of Innovation
by: Agarwal, Shyam, et al.
Published: (2025)
by: Agarwal, Shyam, et al.
Published: (2025)
Detecting algorithmic bias in medical-AI models using trees
by: Smith, Jeffrey, et al.
Published: (2023)
by: Smith, Jeffrey, et al.
Published: (2023)
The Cambridge Law Corpus: A Dataset for Legal AI Research
by: Östling, Andreas, et al.
Published: (2023)
by: Östling, Andreas, et al.
Published: (2023)
AI Evaluation Should Require Standardized Item-Level Data Releases
by: Jiang, Han, et al.
Published: (2026)
by: Jiang, Han, et al.
Published: (2026)
Understanding the Relationship Between Firms' AI Technology Innovation and Consumer Complaints
by: Ma, Yongchao Martin, et al.
Published: (2026)
by: Ma, Yongchao Martin, et al.
Published: (2026)
Generative AI Spotlights the Human Core of Data Science: Implications for Education
by: Taback, Nathan
Published: (2026)
by: Taback, Nathan
Published: (2026)
Leveraging AI for Climate Resilience in Africa: Challenges, Opportunities, and the Need for Collaboration
by: Mbuvha, Rendani, et al.
Published: (2024)
by: Mbuvha, Rendani, et al.
Published: (2024)
Towards Responsible AI in Banking: Addressing Bias for Fair Decision-Making
by: Castelnovo, Alessandro
Published: (2024)
by: Castelnovo, Alessandro
Published: (2024)
A structured regression approach for evaluating model performance across intersectional subgroups
by: Herlihy, Christine, et al.
Published: (2024)
by: Herlihy, Christine, et al.
Published: (2024)
ACRONYM: Augmented degree corrected, Community Reticulated Organized Network Yielding Model
by: Leinwand, Benjamin, et al.
Published: (2024)
by: Leinwand, Benjamin, et al.
Published: (2024)
Vox Populi, Vox AI? Using Language Models to Estimate German Public Opinion
by: von der Heyde, Leah, et al.
Published: (2024)
by: von der Heyde, Leah, et al.
Published: (2024)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Balancing Specialization and Adaptation in a Transforming Scientific Landscape
by: Gautheron, Lucas
Published: (2023)
by: Gautheron, Lucas
Published: (2023)
AI in Work-Based Learning: Understanding the Purposes and Effects of Intelligent Tools Among Student Interns
by: Miranda, John Paul P., et al.
Published: (2026)
by: Miranda, John Paul P., et al.
Published: (2026)
Responsible AI in the Global Context: Maturity Model and Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
The University AI Didn't Replace -- Rethinking Universities in the AI Era
by: Binkowski, Karol P., et al.
Published: (2026)
by: Binkowski, Karol P., et al.
Published: (2026)
Empirical Power Analysis of a Statistical Test to Quantify Gerrymandering
by: Clark, Ranthony A., et al.
Published: (2025)
by: Clark, Ranthony A., et al.
Published: (2025)
Similar Items
-
Dynamic Bayesian Item Response Model with Decomposition (D-BIRD): Modeling Cohort and Individual Learning Over Time
by: Lee, Hansol, et al.
Published: (2025) -
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
by: Salaudeen, Olawale, et al.
Published: (2025) -
Fantastic Bugs and Where to Find Them in AI Benchmarks
by: Truong, Sang, et al.
Published: (2025) -
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
by: Hardy, Michael, et al.
Published: (2026) -
Generative AI Needs Adaptive Governance
by: Reuel, Anka, et al.
Published: (2024)