ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Yizheng, Zeng, Wenjun, Kumaresan, Aditi, Wang, Zi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
por: Hou, Hongru, et al.
Publicado: (2026)
por: Hou, Hongru, et al.
Publicado: (2026)
Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
por: Hahn, Meera, et al.
Publicado: (2024)
por: Hahn, Meera, et al.
Publicado: (2024)
KALAVAI: Predicting When Independent Specialist Fusion Works -- A Quantitative Model for Post-Hoc Cooperative LLM Training
por: Kumaresan, Ramchand
Publicado: (2026)
por: Kumaresan, Ramchand
Publicado: (2026)
ACAR: Adaptive Complexity Routing for Multi-Model Ensembles with Auditable Decision Traces
por: Kumaresan, Ramchand
Publicado: (2026)
por: Kumaresan, Ramchand
Publicado: (2026)
Proactive Depot Discovery: A Generative Framework for Flexible Location-Routing
por: Qu, Site, et al.
Publicado: (2025)
por: Qu, Site, et al.
Publicado: (2025)
ProPACT: A Proactive AI-Driven Adaptive Collaborative Tutor for Pair Programming
por: Golrang, Anahita, et al.
Publicado: (2026)
por: Golrang, Anahita, et al.
Publicado: (2026)
ProAgent: Building Proactive Cooperative Agents with Large Language Models
por: Zhang, Ceyao, et al.
Publicado: (2023)
por: Zhang, Ceyao, et al.
Publicado: (2023)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
por: Wang, Ganghua, et al.
Publicado: (2025)
por: Wang, Ganghua, et al.
Publicado: (2025)
AlphaEval: A Comprehensive and Efficient Evaluation Framework for Formula Alpha Mining
por: Ding, Hongjun, et al.
Publicado: (2025)
por: Ding, Hongjun, et al.
Publicado: (2025)
Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive Assistants
por: Nathani, Deepak, et al.
Publicado: (2026)
por: Nathani, Deepak, et al.
Publicado: (2026)
ULTHO: Ultra-Lightweight yet Efficient Hyperparameter Optimization in Deep Reinforcement Learning
por: Yuan, Mingqi, et al.
Publicado: (2025)
por: Yuan, Mingqi, et al.
Publicado: (2025)
Exploring ChatGPT for Next-generation Information Retrieval: Opportunities and Challenges
por: Huang, Yizheng, et al.
Publicado: (2024)
por: Huang, Yizheng, et al.
Publicado: (2024)
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
por: Li, Mingxuan, et al.
Publicado: (2025)
por: Li, Mingxuan, et al.
Publicado: (2025)
RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation
por: Chen, Shuhao, et al.
Publicado: (2026)
por: Chen, Shuhao, et al.
Publicado: (2026)
MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
por: Zhao, Chenchen, et al.
Publicado: (2025)
por: Zhao, Chenchen, et al.
Publicado: (2025)
Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges
por: Reddy, Chandan K, et al.
Publicado: (2024)
por: Reddy, Chandan K, et al.
Publicado: (2024)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
por: Xu, Weiwei, et al.
Publicado: (2024)
por: Xu, Weiwei, et al.
Publicado: (2024)
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
por: Zhou, Enshen, et al.
Publicado: (2024)
por: Zhou, Enshen, et al.
Publicado: (2024)
Pro-ZD: A Transferable Graph Neural Network Approach for Proactive Zero-Day Threats Mitigation
por: Basta, Nardine, et al.
Publicado: (2026)
por: Basta, Nardine, et al.
Publicado: (2026)
AI and Generative AI for Research Discovery and Summarization
por: Glickman, Mark, et al.
Publicado: (2024)
por: Glickman, Mark, et al.
Publicado: (2024)
AnalogFed: Federated Discovery of Analog Circuit Topologies with Generative AI
por: Li, Qiufeng, et al.
Publicado: (2025)
por: Li, Qiufeng, et al.
Publicado: (2025)
ProCause: Generating Counterfactual Outcomes to Evaluate Prescriptive Process Monitoring Methods
por: De Moor, Jakob, et al.
Publicado: (2025)
por: De Moor, Jakob, et al.
Publicado: (2025)
Symmetric Equilibrium Propagation for Thermodynamic Diffusion Training
por: De, Aditi
Publicado: (2026)
por: De, Aditi
Publicado: (2026)
Thermodynamic Diffusion Inference with Minimal Digital Conditioning
por: De, Aditi
Publicado: (2026)
por: De, Aditi
Publicado: (2026)
Revolutionizing Biomarker Discovery: Leveraging Generative AI for Bio-Knowledge-Embedded Continuous Space Exploration
por: Ying, Wangyang, et al.
Publicado: (2024)
por: Ying, Wangyang, et al.
Publicado: (2024)
Possibility for Proactive Anomaly Detection
por: Jeon, Jinsung, et al.
Publicado: (2025)
por: Jeon, Jinsung, et al.
Publicado: (2025)
PLAN: Proactive Low-Rank Allocation for Continual Learning
por: Wang, Xiequn, et al.
Publicado: (2025)
por: Wang, Xiequn, et al.
Publicado: (2025)
Toward an Evaluation Science for Generative AI Systems
por: Weidinger, Laura, et al.
Publicado: (2025)
por: Weidinger, Laura, et al.
Publicado: (2025)
QualEval: Qualitative Evaluation for Model Improvement
por: Murahari, Vishvak, et al.
Publicado: (2023)
por: Murahari, Vishvak, et al.
Publicado: (2023)
On the Provable Performance Guarantee of Efficient Reasoning Models
por: Zeng, Hao, et al.
Publicado: (2025)
por: Zeng, Hao, et al.
Publicado: (2025)
Evaluation-driven Scaling for Scientific Discovery
por: Ye, Haotian, et al.
Publicado: (2026)
por: Ye, Haotian, et al.
Publicado: (2026)
Graph Federated Learning Based Proactive Content Caching in Edge Computing
por: Wang, Rui
Publicado: (2025)
por: Wang, Rui
Publicado: (2025)
Generating Expressive and Customizable Evals for Timeseries Data Analysis Agents with AgentFuel
por: Maddi, Aadyaa, et al.
Publicado: (2026)
por: Maddi, Aadyaa, et al.
Publicado: (2026)
BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
por: Gandhi, Kanishk, et al.
Publicado: (2025)
por: Gandhi, Kanishk, et al.
Publicado: (2025)
Relative Policy-Transition Optimization for Fast Policy Transfer
por: Xu, Jiawei, et al.
Publicado: (2022)
por: Xu, Jiawei, et al.
Publicado: (2022)
ScholarEval: Research Idea Evaluation Grounded in Literature
por: Moussa, Hanane Nour, et al.
Publicado: (2025)
por: Moussa, Hanane Nour, et al.
Publicado: (2025)
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
por: Grannen, Jennifer, et al.
Publicado: (2025)
por: Grannen, Jennifer, et al.
Publicado: (2025)
Generalized Category Discovery in Federated Graph Learning
por: Yuan, Zhongzheng, et al.
Publicado: (2026)
por: Yuan, Zhongzheng, et al.
Publicado: (2026)
The Use of AI-Robotic Systems for Scientific Discovery
por: Gower, Alexander H., et al.
Publicado: (2024)
por: Gower, Alexander H., et al.
Publicado: (2024)
Model-Based Proactive Cost Generation for Learning Safe Policies Offline with Limited Violation Data
por: Xue, Ruiqi, et al.
Publicado: (2026)
por: Xue, Ruiqi, et al.
Publicado: (2026)
Ejemplares similares
-
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
por: Hou, Hongru, et al.
Publicado: (2026) -
Proactive Agents for Multi-Turn Text-to-Image Generation Under Uncertainty
por: Hahn, Meera, et al.
Publicado: (2024) -
KALAVAI: Predicting When Independent Specialist Fusion Works -- A Quantitative Model for Post-Hoc Cooperative LLM Training
por: Kumaresan, Ramchand
Publicado: (2026) -
ACAR: Adaptive Complexity Routing for Multi-Model Ensembles with Auditable Decision Traces
por: Kumaresan, Ramchand
Publicado: (2026) -
Proactive Depot Discovery: A Generative Framework for Flexible Location-Routing
por: Qu, Site, et al.
Publicado: (2025)