Surrogate-Based Prevalence Measurement for Large-Scale A/B Testing
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Zehao, Paek, Tony, O'Sullivan, Kevin, Dobi, Attila |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Decision Quality Evaluation Framework at Pinterest
por: Tian, Yuqi, et al.
Publicado: (2026)
por: Tian, Yuqi, et al.
Publicado: (2026)
Efficient Prediction of Pass@k Scaling in Large Language Models
por: Kazdan, Joshua, et al.
Publicado: (2025)
por: Kazdan, Joshua, et al.
Publicado: (2025)
SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing
por: Parashar, Anjali, et al.
Publicado: (2026)
por: Parashar, Anjali, et al.
Publicado: (2026)
Conformal Safety Monitoring for Flight Testing: A Case Study in Data-Driven Safety Learning
por: Feldman, Aaron O., et al.
Publicado: (2025)
por: Feldman, Aaron O., et al.
Publicado: (2025)
StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis
por: Song, Xinyi, et al.
Publicado: (2025)
por: Song, Xinyi, et al.
Publicado: (2025)
Performance Evaluation of Large Language Models in Statistical Programming
por: Song, Xinyi, et al.
Publicado: (2025)
por: Song, Xinyi, et al.
Publicado: (2025)
Subnational Geocoding of Global Disasters Using Large Language Models
por: Ronco, Michele, et al.
Publicado: (2025)
por: Ronco, Michele, et al.
Publicado: (2025)
MC-GTA: Metric-Constrained Model-Based Clustering using Goodness-of-fit Tests with Autocorrelations
por: Wang, Zhangyu, et al.
Publicado: (2024)
por: Wang, Zhangyu, et al.
Publicado: (2024)
Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
por: Madden, Emma Rose
Publicado: (2025)
por: Madden, Emma Rose
Publicado: (2025)
Sample-Efficient and Surrogate-Based Design Optimization of Underwater Vehicle Hulls
por: Vardhan, Harsh, et al.
Publicado: (2023)
por: Vardhan, Harsh, et al.
Publicado: (2023)
From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models
por: Zhang, Jiaxin, et al.
Publicado: (2026)
por: Zhang, Jiaxin, et al.
Publicado: (2026)
AI for Handball: predicting and explaining the 2024 Olympic Games tournament with Deep Learning and Large Language Models
por: Felice, Florian
Publicado: (2024)
por: Felice, Florian
Publicado: (2024)
Analyzing the factors that are involved in length of inpatient stay at the hospital for diabetes patients
por: Lam, Jorden, et al.
Publicado: (2024)
por: Lam, Jorden, et al.
Publicado: (2024)
Classification Modeling with RNN-Based, Random Forest, and XGBoost for Imbalanced Data: A Case of Early Crash Detection in ASEAN-5 Stock Markets
por: Siswara, Deri, et al.
Publicado: (2024)
por: Siswara, Deri, et al.
Publicado: (2024)
Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement
por: Sielinski, Ronald
Publicado: (2026)
por: Sielinski, Ronald
Publicado: (2026)
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
por: Oleson, Jon
Publicado: (2024)
por: Oleson, Jon
Publicado: (2024)
E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
por: Sadhuka, Shuvom, et al.
Publicado: (2025)
por: Sadhuka, Shuvom, et al.
Publicado: (2025)
When prompt perturbations break your A/B test: A valid statistical test for generative surveying
por: Helm, Hayden, et al.
Publicado: (2026)
por: Helm, Hayden, et al.
Publicado: (2026)
A Distribution-Free Framework for Rewrite-Based Human-text Detection via Knockoff Filtering
por: Liu, Yi
Publicado: (2026)
por: Liu, Yi
Publicado: (2026)
Scale-Translation Equivariant Network for Oceanic Internal Solitary Wave Localization
por: Wan, Zhang, et al.
Publicado: (2024)
por: Wan, Zhang, et al.
Publicado: (2024)
Cinder: A fast and fair matchmaking system
por: Pal, Saurav
Publicado: (2025)
por: Pal, Saurav
Publicado: (2025)
CERES: A Probabilistic Early Warning System for Acute Food Insecurity
por: Pedersen, Tom Danny S.
Publicado: (2026)
por: Pedersen, Tom Danny S.
Publicado: (2026)
A network analysis of decision strategies of human experts in steel manufacturing
por: Merten, Daniel Christopher, et al.
Publicado: (2021)
por: Merten, Daniel Christopher, et al.
Publicado: (2021)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
por: Luettgau, Lennart, et al.
Publicado: (2025)
por: Luettgau, Lennart, et al.
Publicado: (2025)
Process-Aware Analysis of Treatment Paths in Heart Failure Patients: A Case Study
por: Beyel, Harry H., et al.
Publicado: (2024)
por: Beyel, Harry H., et al.
Publicado: (2024)
TCKAN:A Novel Integrated Network Model for Predicting Mortality Risk in Sepsis Patients
por: Dong, Fanglin
Publicado: (2024)
por: Dong, Fanglin
Publicado: (2024)
Eligibility-Aware Evidence Synthesis: An Agentic Framework for Clinical Trial Meta-Analysis
por: Zhao, Yao, et al.
Publicado: (2026)
por: Zhao, Yao, et al.
Publicado: (2026)
A survey of using EHR as real-world evidence for discovering and validating new drug indications
por: Talukdar, Nabasmita, et al.
Publicado: (2025)
por: Talukdar, Nabasmita, et al.
Publicado: (2025)
A Regression Mixture Model to understand the effect of the Covid-19 pandemic on Public Transport Ridership
por: Moreau, Hugues, et al.
Publicado: (2024)
por: Moreau, Hugues, et al.
Publicado: (2024)
A Statistical Theory of Regularization-Based Continual Learning
por: Zhao, Xuyang, et al.
Publicado: (2024)
por: Zhao, Xuyang, et al.
Publicado: (2024)
Prediction of Delirium Risk in Mild Cognitive Impairment Using Time-Series data, Machine Learning and Comorbidity Patterns -- A Retrospective Study
por: Ramamoorthy, Santhakumar, et al.
Publicado: (2025)
por: Ramamoorthy, Santhakumar, et al.
Publicado: (2025)
FactsR: A Safer Method for Producing High Quality Healthcare Documentation
por: Hansen, Victor Petrén Bach, et al.
Publicado: (2025)
por: Hansen, Victor Petrén Bach, et al.
Publicado: (2025)
Personalization of Large Foundation Models for Health Interventions
por: Konigorski, Stefan, et al.
Publicado: (2026)
por: Konigorski, Stefan, et al.
Publicado: (2026)
Acquiring Better Load Estimates by Combining Anomaly and Change Point Detection in Power Grid Time-series Measurements
por: Bouman, Roel, et al.
Publicado: (2024)
por: Bouman, Roel, et al.
Publicado: (2024)
Confidence Adjusted Surprise Measure for Active Resourceful Trials (CA-SMART): A Data-driven Active Learning Framework for Accelerating Material Discovery under Resource Constraints
por: Raihan, Ahmed Shoyeb, et al.
Publicado: (2025)
por: Raihan, Ahmed Shoyeb, et al.
Publicado: (2025)
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
por: Xu, Zhiyu, et al.
Publicado: (2025)
por: Xu, Zhiyu, et al.
Publicado: (2025)
TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models
por: Devunuri, Saipraneeth, et al.
Publicado: (2024)
por: Devunuri, Saipraneeth, et al.
Publicado: (2024)
Can-SAVE: Deploying Low-Cost and Population-Scale Cancer Screening via Survival Analysis Variables and EHR
por: Philonenko, Petr, et al.
Publicado: (2023)
por: Philonenko, Petr, et al.
Publicado: (2023)
Automated Vehicles at Unsignalized Intersections: Safety and Efficiency Implications of Mixed Human and Automated Traffic
por: Rahmani, Saeed, et al.
Publicado: (2024)
por: Rahmani, Saeed, et al.
Publicado: (2024)
The Evolution of Probabilistic Price Forecasting Techniques: A Review of the Day-Ahead, Intra-Day, and Balancing Markets
por: O'Connor, Ciaran, et al.
Publicado: (2025)
por: O'Connor, Ciaran, et al.
Publicado: (2025)
Ejemplares similares
-
Decision Quality Evaluation Framework at Pinterest
por: Tian, Yuqi, et al.
Publicado: (2026) -
Efficient Prediction of Pass@k Scaling in Large Language Models
por: Kazdan, Joshua, et al.
Publicado: (2025) -
SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing
por: Parashar, Anjali, et al.
Publicado: (2026) -
Conformal Safety Monitoring for Flight Testing: A Case Study in Data-Driven Safety Learning
por: Feldman, Aaron O., et al.
Publicado: (2025) -
StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis
por: Song, Xinyi, et al.
Publicado: (2025)