Decision Quality Evaluation Framework at Pinterest
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Yuqi, Paine, Robert, Dobi, Attila, O'Sullivan, Kevin, Manickavasagam, Aravindh, Farooq, Faisal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Surrogate-Based Prevalence Measurement for Large-Scale A/B Testing
by: Xu, Zehao, et al.
Published: (2026)
by: Xu, Zehao, et al.
Published: (2026)
Measuring the Prevalence of Policy Violating Content with ML Assisted Sampling and LLM Labeling
by: Dobi, Attila, et al.
Published: (2026)
by: Dobi, Attila, et al.
Published: (2026)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025)
by: Luettgau, Lennart, et al.
Published: (2025)
Data-Driven Bayesian Network Models of Hurricane Evacuation Decision Making
by: Wang, Hui Sophie, et al.
Published: (2023)
by: Wang, Hui Sophie, et al.
Published: (2023)
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
by: Xu, Zhiyu, et al.
Published: (2025)
by: Xu, Zhiyu, et al.
Published: (2025)
FactsR: A Safer Method for Producing High Quality Healthcare Documentation
by: Hansen, Victor Petrén Bach, et al.
Published: (2025)
by: Hansen, Victor Petrén Bach, et al.
Published: (2025)
AI-Assisted Decision-Making for Clinical Assessment of Auto-Segmented Contour Quality
by: Wang, Biling, et al.
Published: (2025)
by: Wang, Biling, et al.
Published: (2025)
Performance Evaluation of Large Language Models in Statistical Programming
by: Song, Xinyi, et al.
Published: (2025)
by: Song, Xinyi, et al.
Published: (2025)
Embracing Ambiguity: Bayesian Nonparametrics and Stakeholder Participation for Ambiguity-Aware Safety Evaluation
by: Long, Yanan
Published: (2025)
by: Long, Yanan
Published: (2025)
StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis
by: Song, Xinyi, et al.
Published: (2025)
by: Song, Xinyi, et al.
Published: (2025)
Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
by: Madden, Emma Rose
Published: (2025)
by: Madden, Emma Rose
Published: (2025)
Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction
by: Vishwakarma, Harit, et al.
Published: (2024)
by: Vishwakarma, Harit, et al.
Published: (2024)
Sentiment Analysis Based on RoBERTa for Amazon Review: An Empirical Study on Decision Making
by: Guo, Xinli
Published: (2024)
by: Guo, Xinli
Published: (2024)
Efficient Prediction of Pass@k Scaling in Large Language Models
by: Kazdan, Joshua, et al.
Published: (2025)
by: Kazdan, Joshua, et al.
Published: (2025)
Eligibility-Aware Evidence Synthesis: An Agentic Framework for Clinical Trial Meta-Analysis
by: Zhao, Yao, et al.
Published: (2026)
by: Zhao, Yao, et al.
Published: (2026)
A Distribution-Free Framework for Rewrite-Based Human-text Detection via Knockoff Filtering
by: Liu, Yi
Published: (2026)
by: Liu, Yi
Published: (2026)
Multi-spatial Multi-temporal Air Quality Forecasting with Integrated Monitoring and Reanalysis Data
by: Hu, Yuxiao, et al.
Published: (2023)
by: Hu, Yuxiao, et al.
Published: (2023)
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
by: Oleson, Jon
Published: (2024)
by: Oleson, Jon
Published: (2024)
Binary Gaussian Copula Synthesis: A Novel Data Augmentation Technique to Advance ML-based Clinical Decision Support Systems for Early Prediction of Dialysis Among CKD Patients
by: Khosravi, Hamed, et al.
Published: (2024)
by: Khosravi, Hamed, et al.
Published: (2024)
ACT-Tensor: Tensor Completion Framework for Financial Dataset Imputation
by: Mo, Junyi, et al.
Published: (2025)
by: Mo, Junyi, et al.
Published: (2025)
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
by: Hardy, Michael
Published: (2024)
by: Hardy, Michael
Published: (2024)
AI-Assisted Conversational Interviewing: Effects on Data Quality and Respondent Experience
by: Barari, Soubhik, et al.
Published: (2025)
by: Barari, Soubhik, et al.
Published: (2025)
A Cybersecurity Risk Analysis Framework for Systems with Artificial Intelligence Components
by: Camacho, Jose Manuel, et al.
Published: (2024)
by: Camacho, Jose Manuel, et al.
Published: (2024)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
by: Xu, Yang, et al.
Published: (2026)
by: Xu, Yang, et al.
Published: (2026)
Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement
by: Sielinski, Ronald
Published: (2026)
by: Sielinski, Ronald
Published: (2026)
What If They Took the Shot? A Hierarchical Bayesian Framework for Counterfactual Expected Goals
by: Mahmudlu, Mikayil, et al.
Published: (2025)
by: Mahmudlu, Mikayil, et al.
Published: (2025)
Temporal Subtyping of Alzheimer's Disease Using Medical Conditions Preceding Alzheimer's Disease Onset in Electronic Health Records
by: He, Zhe, et al.
Published: (2022)
by: He, Zhe, et al.
Published: (2022)
Collective Reasoning Among LLMs: A Framework for Answer Validation Without Ground Truth
by: Davoudi, Seyed Pouyan Mousavi, et al.
Published: (2025)
by: Davoudi, Seyed Pouyan Mousavi, et al.
Published: (2025)
A More Realistic Evaluation of Cross-Frequency Transfer Learning and Foundation Forecasting Models
by: Olivares, Kin G., et al.
Published: (2025)
by: Olivares, Kin G., et al.
Published: (2025)
SureMap: Simultaneous Mean Estimation for Single-Task and Multi-Task Disaggregated Evaluation
by: Khodak, Mikhail, et al.
Published: (2024)
by: Khodak, Mikhail, et al.
Published: (2024)
Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys
by: Ye, Zikun, et al.
Published: (2026)
by: Ye, Zikun, et al.
Published: (2026)
CERES: A Probabilistic Early Warning System for Acute Food Insecurity
by: Pedersen, Tom Danny S.
Published: (2026)
by: Pedersen, Tom Danny S.
Published: (2026)
SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing
by: Parashar, Anjali, et al.
Published: (2026)
by: Parashar, Anjali, et al.
Published: (2026)
From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
by: Hardy, Michael, et al.
Published: (2026)
by: Hardy, Michael, et al.
Published: (2026)
On the Mechanistic Interpretability of Neural Networks for Causality in Bio-statistics
by: Conan, Jean-Baptiste A.
Published: (2025)
by: Conan, Jean-Baptiste A.
Published: (2025)
Quantitative Technology Forecasting: a Review of Trend Extrapolation Methods
by: Tsai, Peng-Hung, et al.
Published: (2024)
by: Tsai, Peng-Hung, et al.
Published: (2024)
Decade-long Emission Forecasting with an Ensemble Model in Taiwan
by: Hung, Gordon, et al.
Published: (2025)
by: Hung, Gordon, et al.
Published: (2025)
ChatGPT and post-test probability
by: Weisenthal, Samuel J.
Published: (2023)
by: Weisenthal, Samuel J.
Published: (2023)
Process-Aware Analysis of Treatment Paths in Heart Failure Patients: A Case Study
by: Beyel, Harry H., et al.
Published: (2024)
by: Beyel, Harry H., et al.
Published: (2024)
Similar Items
-
Surrogate-Based Prevalence Measurement for Large-Scale A/B Testing
by: Xu, Zehao, et al.
Published: (2026) -
Measuring the Prevalence of Policy Violating Content with ML Assisted Sampling and LLM Labeling
by: Dobi, Attila, et al.
Published: (2026) -
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
by: Luettgau, Lennart, et al.
Published: (2025) -
Data-Driven Bayesian Network Models of Hurricane Evacuation Decision Making
by: Wang, Hui Sophie, et al.
Published: (2023) -
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
by: Xu, Zhiyu, et al.
Published: (2025)