AutoEval Done Right: Using Synthetic Data for Model Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Boyeau, Pierre, Angelopoulos, Anastasios N., Yosef, Nir, Malik, Jitendra, Jordan, Michael I. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
von: Park, Sangwoo, et al.
Veröffentlicht: (2025)
von: Park, Sangwoo, et al.
Veröffentlicht: (2025)
Conformal Decision Theory: Safe Autonomous Decisions from Imperfect Predictions
von: Lekeufack, Jordan, et al.
Veröffentlicht: (2023)
von: Lekeufack, Jordan, et al.
Veröffentlicht: (2023)
Private Prediction Sets
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2021)
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2021)
Data-Adaptive Tradeoffs among Multiple Risks in Distribution-Free Prediction
von: Nguyen, Drew T., et al.
Veröffentlicht: (2024)
von: Nguyen, Drew T., et al.
Veröffentlicht: (2024)
AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
Conformal Risk Control for Non-Monotonic Losses
von: Angelopoulos, Anastasios N.
Veröffentlicht: (2026)
von: Angelopoulos, Anastasios N.
Veröffentlicht: (2026)
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2025)
Using Large Language Models to Suggest Informative Prior Distributions in Bayesian Statistics
von: Riegler, Michael A., et al.
Veröffentlicht: (2025)
von: Riegler, Michael A., et al.
Veröffentlicht: (2025)
Conformal Prediction Under Feedback Covariate Shift for Biomolecular Design
von: Fannjiang, Clara, et al.
Veröffentlicht: (2022)
von: Fannjiang, Clara, et al.
Veröffentlicht: (2022)
GLM Inference with AI-Generated Synthetic Data Using Misspecified Linear Regression
von: Keret, Nir, et al.
Veröffentlicht: (2025)
von: Keret, Nir, et al.
Veröffentlicht: (2025)
Conformal Risk Control
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2022)
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2022)
Batch Speculative Decoding Done Right
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
Label Noise Robustness of Conformal Prediction
von: Einbinder, Bat-Sheva, et al.
Veröffentlicht: (2022)
von: Einbinder, Bat-Sheva, et al.
Veröffentlicht: (2022)
Online conformal prediction with decaying step sizes
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2024)
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2024)
PPI++: Efficient Prediction-Powered Inference
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2023)
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2023)
Evaluating Interventional Reasoning Capabilities of Large Language Models
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
Human-AI Co-design for Clinical Prediction Models
von: Feng, Jean, et al.
Veröffentlicht: (2026)
von: Feng, Jean, et al.
Veröffentlicht: (2026)
Theoretical Foundations of Conformal Prediction
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2024)
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2024)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
A Causal Lens for Evaluating Faithfulness Metrics
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
RCT Rejection Sampling for Causal Estimation Evaluation
von: Keith, Katherine A., et al.
Veröffentlicht: (2023)
von: Keith, Katherine A., et al.
Veröffentlicht: (2023)
A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
Banking Done Right: Redefining Retail Banking with Language-Centric AI
von: Chua, Xin Jie, et al.
Veröffentlicht: (2025)
von: Chua, Xin Jie, et al.
Veröffentlicht: (2025)
Synthetic Data, Information, and Prior Knowledge: Why Synthetic Data Augmentation to Boost Sample Doesn't Work for Statistical Inference
von: Dale, Reid, et al.
Veröffentlicht: (2026)
von: Dale, Reid, et al.
Veröffentlicht: (2026)
Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content
von: Bednarczyk, Lydie, et al.
Veröffentlicht: (2026)
von: Bednarczyk, Lydie, et al.
Veröffentlicht: (2026)
Segmenting Human-LLM Co-authored Text via Change Point Detection
von: Li, Mengchu, et al.
Veröffentlicht: (2026)
von: Li, Mengchu, et al.
Veröffentlicht: (2026)
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
von: Cai, Hengrui, et al.
Veröffentlicht: (2023)
von: Cai, Hengrui, et al.
Veröffentlicht: (2023)
CLEAR: Can Language Models Really Understand Causal Graphs?
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
Trustworthy Evaluation of Robotic Manipulation: A New Benchmark and AutoEval Methods
von: Liu, Mengyuan, et al.
Veröffentlicht: (2026)
von: Liu, Mengyuan, et al.
Veröffentlicht: (2026)
Conformal Prediction Sets with Improved Conditional Coverage using Trust Scores
von: Kaur, Jivat Neet, et al.
Veröffentlicht: (2025)
von: Kaur, Jivat Neet, et al.
Veröffentlicht: (2025)
LLMs are Overconfident: Evaluating Confidence Interval Calibration with FermiEval
von: Epstein, Elliot L., et al.
Veröffentlicht: (2025)
von: Epstein, Elliot L., et al.
Veröffentlicht: (2025)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
Super-Level-Set Regression: Conditional Quantiles via Volume Minimization
von: Braun, Sacha, et al.
Veröffentlicht: (2026)
von: Braun, Sacha, et al.
Veröffentlicht: (2026)
Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities
von: Xu, Yichen
Veröffentlicht: (2026)
von: Xu, Yichen
Veröffentlicht: (2026)
Language Models as Causal Effect Generators
von: Bynum, Lucius E. J., et al.
Veröffentlicht: (2024)
von: Bynum, Lucius E. J., et al.
Veröffentlicht: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
von: Hobelsberger, Christian, et al.
Veröffentlicht: (2025)
von: Hobelsberger, Christian, et al.
Veröffentlicht: (2025)
Grounding Synthetic Data Evaluations of Language Models in Unsupervised Document Corpora
von: Majurski, Michael, et al.
Veröffentlicht: (2025)
von: Majurski, Michael, et al.
Veröffentlicht: (2025)
Are We Done with MMLU?
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
von: Park, Sangwoo, et al.
Veröffentlicht: (2025) -
Conformal Decision Theory: Safe Autonomous Decisions from Imperfect Predictions
von: Lekeufack, Jordan, et al.
Veröffentlicht: (2023) -
Private Prediction Sets
von: Angelopoulos, Anastasios N., et al.
Veröffentlicht: (2021) -
Data-Adaptive Tradeoffs among Multiple Risks in Distribution-Free Prediction
von: Nguyen, Drew T., et al.
Veröffentlicht: (2024) -
AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)