CIRCLE: A Framework for Evaluating AI from a Real-World Lens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schwartz, Reva, Westling, Carina, Briggs, Morgan, Fadaee, Marzieh, Nejadgholi, Isar, Holmes, Matthew, Rashid, Fariza, Carlyle, Maya, Taïk, Afaf, Wilson, Kyra, Douglas, Peter, Skeadas, Theodora, Waters, Gabriella, Chowdhury, Rumman, Lacerda, Thiago |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
von: Schwartz, Reva, et al.
Veröffentlicht: (2025)
von: Schwartz, Reva, et al.
Veröffentlicht: (2025)
Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Maker's Dilemma
von: Schwartz, Reva, et al.
Veröffentlicht: (2026)
von: Schwartz, Reva, et al.
Veröffentlicht: (2026)
Making AI Evaluation Deployment Relevant Through Context Specification
von: Holmes, Matthew, et al.
Veröffentlicht: (2026)
von: Holmes, Matthew, et al.
Veröffentlicht: (2026)
Ask What Your Country Can Do For You: Towards a Public Red Teaming Model
von: Kennedy, Wm. Matthew, et al.
Veröffentlicht: (2025)
von: Kennedy, Wm. Matthew, et al.
Veröffentlicht: (2025)
WildFireCan-MMD: A Multimodal Dataset for Classification of User-Generated Content During Wildfires in Canada
von: Sherritt, Braeden, et al.
Veröffentlicht: (2025)
von: Sherritt, Braeden, et al.
Veröffentlicht: (2025)
Automatic WordNet Construction Using Markov Chain Monte Carlo
von: Marzieh Fadaee
Veröffentlicht: (2013)
von: Marzieh Fadaee
Veröffentlicht: (2013)
Social and Ethical Risks Posed by General-Purpose LLMs for Settling Newcomers in Canada
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2024)
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2024)
Systemizing Multiplicity: The Curious Case of Arbitrariness in Machine Learning
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2025)
von: Ganesh, Prakhar, et al.
Veröffentlicht: (2025)
Fairness in Federated Learning: Fairness for Whom?
von: Taik, Afaf, et al.
Veröffentlicht: (2025)
von: Taik, Afaf, et al.
Veröffentlicht: (2025)
Multilingual Hallucination Gaps in Large Language Models
von: Chataigner, Cléa, et al.
Veröffentlicht: (2024)
von: Chataigner, Cléa, et al.
Veröffentlicht: (2024)
Promoting Fair Vaccination Strategies Through Influence Maximization: A Case Study on COVID-19 Spread
von: Neophytou, Nicola, et al.
Veröffentlicht: (2024)
von: Neophytou, Nicola, et al.
Veröffentlicht: (2024)
Differentially Private Clustered Federated Learning
von: Malekmohammadi, Saber, et al.
Veröffentlicht: (2024)
von: Malekmohammadi, Saber, et al.
Veröffentlicht: (2024)
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
von: Dawkins, Hillary, et al.
Veröffentlicht: (2024)
von: Dawkins, Hillary, et al.
Veröffentlicht: (2024)
Gender-Neutral Machine Translation Strategies in Practice
von: Dawkins, Hillary, et al.
Veröffentlicht: (2025)
von: Dawkins, Hillary, et al.
Veröffentlicht: (2025)
Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
von: Ghanadian, Hamideh, et al.
Veröffentlicht: (2024)
von: Ghanadian, Hamideh, et al.
Veröffentlicht: (2024)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2026)
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2026)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
von: Vishnubhotla, Krishnapriya, et al.
Veröffentlicht: (2026)
von: Vishnubhotla, Krishnapriya, et al.
Veröffentlicht: (2026)
Human-Centered AI Applications for Canada's Immigration Settlement Sector
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2024)
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2024)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2025)
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2025)
From Perceived Effectiveness to Measured Impact: Identity-Aware Evaluation of Automated Counter-Stereotypes
von: Kiritchenko, Svetlana, et al.
Veröffentlicht: (2025)
von: Kiritchenko, Svetlana, et al.
Veröffentlicht: (2025)
Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
von: Fraser, Kathleen C., et al.
Veröffentlicht: (2025)
von: Fraser, Kathleen C., et al.
Veröffentlicht: (2025)
The crime of being poor
von: Curto, Georgina, et al.
Veröffentlicht: (2023)
von: Curto, Georgina, et al.
Veröffentlicht: (2023)
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
von: Dawkins, Hillary, et al.
Veröffentlicht: (2024)
von: Dawkins, Hillary, et al.
Veröffentlicht: (2024)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2024)
von: Nejadgholi, Isar, et al.
Veröffentlicht: (2024)
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
von: Guo, Rongchen, et al.
Veröffentlicht: (2025)
von: Guo, Rongchen, et al.
Veröffentlicht: (2025)
Fairness Incentives in Response to Unfair Dynamic Pricing
von: Thibodeau, Jesse, et al.
Veröffentlicht: (2024)
von: Thibodeau, Jesse, et al.
Veröffentlicht: (2024)
Balancing Profit and Fairness in Risk-Based Pricing Markets
von: Thibodeau, Jesse, et al.
Veröffentlicht: (2025)
von: Thibodeau, Jesse, et al.
Veröffentlicht: (2025)
Uncanny or Not? Perceptions of AI-Generated Faces in Autism
von: Waters, Gabriella
Veröffentlicht: (2025)
von: Waters, Gabriella
Veröffentlicht: (2025)
Testing, Evaluation, Verification and Validation (TEVV) of Digital Twins: A Comprehensive Framework
von: Waters, Gabriella
Veröffentlicht: (2025)
von: Waters, Gabriella
Veröffentlicht: (2025)
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
von: Guo, Rongchen, et al.
Veröffentlicht: (2024)
von: Guo, Rongchen, et al.
Veröffentlicht: (2024)
PRAGMATIK ADABIYOTSHUNOSLIK FAN SOHASI SIFATIDA: MUTOLAA FENOMENOLOGIYASI MISOLIDA
von: Maxmudjonova Fariza
Veröffentlicht: (2025)
von: Maxmudjonova Fariza
Veröffentlicht: (2025)
Proof of Concept for Mammography Classification with Enhanced Compactness and Separability Modules
von: Dahes, Fariza
Veröffentlicht: (2025)
von: Dahes, Fariza
Veröffentlicht: (2025)
LLMs are One-Shot URL Classifiers and Explainers
von: Rashid, Fariza, et al.
Veröffentlicht: (2024)
von: Rashid, Fariza, et al.
Veröffentlicht: (2024)
Eliciting Least-to-Most Reasoning for Phishing URL Detection
von: Trikilis, Holly, et al.
Veröffentlicht: (2026)
von: Trikilis, Holly, et al.
Veröffentlicht: (2026)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
von: Shimabucoro, Luisa, et al.
Veröffentlicht: (2025)
von: Shimabucoro, Luisa, et al.
Veröffentlicht: (2025)
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
von: Yu, Simon, et al.
Veröffentlicht: (2024)
von: Yu, Simon, et al.
Veröffentlicht: (2024)
Manifold-Matching Autoencoders
von: Cheret, Laurent, et al.
Veröffentlicht: (2026)
von: Cheret, Laurent, et al.
Veröffentlicht: (2026)
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
von: Curto, Georgina, et al.
Veröffentlicht: (2025)
von: Curto, Georgina, et al.
Veröffentlicht: (2025)
Nonparametric Sensitivity Analysis for Unobserved Confounding with Survival Outcomes
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
von: Schwartz, Reva, et al.
Veröffentlicht: (2025) -
Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Maker's Dilemma
von: Schwartz, Reva, et al.
Veröffentlicht: (2026) -
Making AI Evaluation Deployment Relevant Through Context Specification
von: Holmes, Matthew, et al.
Veröffentlicht: (2026) -
Ask What Your Country Can Do For You: Towards a Public Red Teaming Model
von: Kennedy, Wm. Matthew, et al.
Veröffentlicht: (2025) -
WildFireCan-MMD: A Multimodal Dataset for Classification of User-Generated Content During Wildfires in Canada
von: Sherritt, Braeden, et al.
Veröffentlicht: (2025)