Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
Fuente:
arXiv
Saved in:
| Main Authors: | Reuel, Anka, Ghosh, Avijit, Chim, Jenny, Tran, Andrew, Long, Yanan, Mickel, Jennifer, Gohar, Usman, Yadav, Srishti, Ammanamanchi, Pawan Sasanka, Allaham, Mowafak, Rahmani, Hossein A., Akhtar, Mubashara, Friedrich, Felix, Scholz, Robert, Riegler, Michael Alexander, Batzner, Jan, Habba, Eliya, Saxena, Arushi, Kornilova, Anastassia, Wei, Kevin, Soni, Prajna, Mathew, Yohan, Klyman, Kevin, Sania, Jeba, Sahoo, Subramanyam, Bruvik, Olivia Beyer, Sadeghi, Pouya, Goswami, Sujata, Wang, Angelina, Jernite, Yacine, Talat, Zeerak, Biderman, Stella, Kochenderfer, Mykel, Koyejo, Sanmi, Solaiman, Irene |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
by: Allaham, Mowafak, et al.
Published: (2024)
by: Allaham, Mowafak, et al.
Published: (2024)
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
by: Akhtar, Mubashara, et al.
Published: (2026)
by: Akhtar, Mubashara, et al.
Published: (2026)
Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources
by: Allaham, Mowafak, et al.
Published: (2026)
by: Allaham, Mowafak, et al.
Published: (2026)
Towards Leveraging News Media to Support Impact Assessment of AI Technologies
by: Allaham, Mowafak, et al.
Published: (2024)
by: Allaham, Mowafak, et al.
Published: (2024)
Global Perspectives of AI Risks and Harms: Analyzing the Negative Impacts of AI Technologies as Prioritized by News Media
by: Allaham, Mowafak, et al.
Published: (2025)
by: Allaham, Mowafak, et al.
Published: (2025)
Informing AI Risk Assessment with News Media: Analyzing National and Political Variation in the Coverage of AI Risks
by: Allaham, Mowafak, et al.
Published: (2025)
by: Allaham, Mowafak, et al.
Published: (2025)
Emergence WebVoyager: Toward Consistent and Transparent Evaluation of (Web) Agents in The Wild
by: Akkil, Deepak, et al.
Published: (2026)
by: Akkil, Deepak, et al.
Published: (2026)
Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change
by: Allaham, Mowafak, et al.
Published: (2025)
by: Allaham, Mowafak, et al.
Published: (2025)
Audit Cards: Contextualizing AI Evaluations
by: Staufer, Leon, et al.
Published: (2025)
by: Staufer, Leon, et al.
Published: (2025)
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
by: Lamparth, Max, et al.
Published: (2023)
by: Lamparth, Max, et al.
Published: (2023)
Fairness in Reinforcement Learning: A Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
by: Lior, Gili, et al.
Published: (2025)
by: Lior, Gili, et al.
Published: (2025)
Acceptable Use Policies for Foundation Models
by: Klyman, Kevin
Published: (2024)
by: Klyman, Kevin
Published: (2024)
Generative AI Needs Adaptive Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
JSON Whisperer: Efficient JSON Editing with LLMs
by: Duanis, Sarel, et al.
Published: (2025)
by: Duanis, Sarel, et al.
Published: (2025)
Beyond Benchmarks: On The False Promise of AI Regulation
by: Stanovsky, Gabriel, et al.
Published: (2025)
by: Stanovsky, Gabriel, et al.
Published: (2025)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
by: Itzhak, Itay, et al.
Published: (2026)
by: Itzhak, Itay, et al.
Published: (2026)
From Symptoms to Systems: An Expert-Guided Approach to Understanding Risks of Generative AI for Eating Disorders
by: Winecoff, Amy, et al.
Published: (2025)
by: Winecoff, Amy, et al.
Published: (2025)
SpecEval: Evaluating Model Adherence to Behavior Specifications
by: Ahmed, Ahmed, et al.
Published: (2025)
by: Ahmed, Ahmed, et al.
Published: (2025)
Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
by: Wei, Kevin L., et al.
Published: (2025)
by: Wei, Kevin L., et al.
Published: (2025)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
by: Habba, Eliya, et al.
Published: (2025)
by: Habba, Eliya, et al.
Published: (2025)
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
by: Levy, Shahar, et al.
Published: (2026)
by: Levy, Shahar, et al.
Published: (2026)
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking
by: Akhtar, Mubashara, et al.
Published: (2024)
by: Akhtar, Mubashara, et al.
Published: (2024)
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
by: Haupt, Andreas, et al.
Published: (2026)
by: Haupt, Andreas, et al.
Published: (2026)
Do AI Companies Make Good on Voluntary Commitments to the White House?
by: Wang, Jennifer, et al.
Published: (2025)
by: Wang, Jennifer, et al.
Published: (2025)
Spectral statistics and energy-gap scaling in $k-$local spin Hamiltonians
by: Dowarah, Sasanka
Published: (2025)
by: Dowarah, Sasanka
Published: (2025)
Beyond Release: Access Considerations for Generative AI Systems
by: Solaiman, Irene, et al.
Published: (2025)
by: Solaiman, Irene, et al.
Published: (2025)
Evaluating the Social Impact of Generative AI Systems in Systems and Society
by: Solaiman, Irene, et al.
Published: (2023)
by: Solaiman, Irene, et al.
Published: (2023)
Two New Species of Elaphoglossum (Elaphoglossaceae) from Amazonas, Venezuela
by: Mickel, John T.
Published: (1990)
by: Mickel, John T.
Published: (1990)
Racial/Ethnic Categories in AI and Algorithmic Fairness: Why They Matter and What They Represent
by: Mickel, Jennifer
Published: (2024)
by: Mickel, Jennifer
Published: (2024)
More of the Same: Persistent Representational Harms Under Increased Representation
by: Mickel, Jennifer, et al.
Published: (2025)
by: Mickel, Jennifer, et al.
Published: (2025)
New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses
by: McDuff, Daniel, et al.
Published: (2025)
by: McDuff, Daniel, et al.
Published: (2025)
Recourse, Repair, Reparation, & Prevention: A Stakeholder Analysis of AI Supply Chains
by: Hopkins, Aspen K., et al.
Published: (2025)
by: Hopkins, Aspen K., et al.
Published: (2025)
Comparing Apples to Oranges: A Taxonomy for Navigating the Global Landscape of AI Regulation
by: Alanoca, Sacha, et al.
Published: (2025)
by: Alanoca, Sacha, et al.
Published: (2025)
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
by: Rivera, Juan-Pablo, et al.
Published: (2024)
by: Rivera, Juan-Pablo, et al.
Published: (2024)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Lessons from the Trenches on Reproducible Evaluation of Language Models
by: Biderman, Stella, et al.
Published: (2024)
by: Biderman, Stella, et al.
Published: (2024)
Similar Items
-
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
by: Allaham, Mowafak, et al.
Published: (2024) -
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
by: Akhtar, Mubashara, et al.
Published: (2026) -
Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources
by: Allaham, Mowafak, et al.
Published: (2026) -
Towards Leveraging News Media to Support Impact Assessment of AI Technologies
by: Allaham, Mowafak, et al.
Published: (2024) -
Global Perspectives of AI Risks and Harms: Analyzing the Negative Impacts of AI Technologies as Prioritized by News Media
by: Allaham, Mowafak, et al.
Published: (2025)