Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Ganghua, Chen, Zhaorun, Li, Bo, Xu, Haifeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Certifying Knowledge Comprehension in LLMs
por: Chaudhary, Isha, et al.
Publicado: (2024)
por: Chaudhary, Isha, et al.
Publicado: (2024)
Cost-Effective Hallucination Detection for LLMs
por: Valentin, Simon, et al.
Publicado: (2024)
por: Valentin, Simon, et al.
Publicado: (2024)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
por: Sansford, Hannah, et al.
Publicado: (2024)
por: Sansford, Hannah, et al.
Publicado: (2024)
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
por: Li, Mingxuan, et al.
Publicado: (2025)
por: Li, Mingxuan, et al.
Publicado: (2025)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
por: Chung, Tsz Ting, et al.
Publicado: (2025)
por: Chung, Tsz Ting, et al.
Publicado: (2025)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
por: Hu, Haiquan, et al.
Publicado: (2025)
por: Hu, Haiquan, et al.
Publicado: (2025)
QualEval: Qualitative Evaluation for Model Improvement
por: Murahari, Vishvak, et al.
Publicado: (2023)
por: Murahari, Vishvak, et al.
Publicado: (2023)
Anyprefer: An Agentic Framework for Preference Data Synthesis
por: Zhou, Yiyang, et al.
Publicado: (2025)
por: Zhou, Yiyang, et al.
Publicado: (2025)
ScholarEval: Research Idea Evaluation Grounded in Literature
por: Moussa, Hanane Nour, et al.
Publicado: (2025)
por: Moussa, Hanane Nour, et al.
Publicado: (2025)
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
por: Liu, Minqian, et al.
Publicado: (2023)
por: Liu, Minqian, et al.
Publicado: (2023)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
por: Wang, Guangtao, et al.
Publicado: (2025)
por: Wang, Guangtao, et al.
Publicado: (2025)
Measuring all the noises of LLM Evals
por: Wang, Sida
Publicado: (2025)
por: Wang, Sida
Publicado: (2025)
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
por: Siu, Vincent, et al.
Publicado: (2025)
por: Siu, Vincent, et al.
Publicado: (2025)
EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
por: Nie, Allen, et al.
Publicado: (2024)
por: Nie, Allen, et al.
Publicado: (2024)
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
por: Peng, Bo, et al.
Publicado: (2026)
por: Peng, Bo, et al.
Publicado: (2026)
HumanEval on Latest GPT Models -- 2024
por: Li, Daniel, et al.
Publicado: (2024)
por: Li, Daniel, et al.
Publicado: (2024)
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
por: Dubois, Yann, et al.
Publicado: (2024)
por: Dubois, Yann, et al.
Publicado: (2024)
Sample-Efficient Alignment for LLMs
por: Liu, Zichen, et al.
Publicado: (2024)
por: Liu, Zichen, et al.
Publicado: (2024)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
por: Cao, Boxi, et al.
Publicado: (2024)
por: Cao, Boxi, et al.
Publicado: (2024)
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking
por: Xu, Qinwu, et al.
Publicado: (2026)
por: Xu, Qinwu, et al.
Publicado: (2026)
xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
por: Qian, Cheng, et al.
Publicado: (2025)
por: Qian, Cheng, et al.
Publicado: (2025)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
por: He, Yuanqin, et al.
Publicado: (2024)
por: He, Yuanqin, et al.
Publicado: (2024)
A Unified Framework with Novel Metrics for Evaluating the Effectiveness of XAI Techniques in LLMs
por: Mersha, Melkamu Abay, et al.
Publicado: (2025)
por: Mersha, Melkamu Abay, et al.
Publicado: (2025)
Efficient multi-prompt evaluation of LLMs
por: Polo, Felipe Maia, et al.
Publicado: (2024)
por: Polo, Felipe Maia, et al.
Publicado: (2024)
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
por: Zeng, Qiuhai, et al.
Publicado: (2025)
por: Zeng, Qiuhai, et al.
Publicado: (2025)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
por: Boyeau, Pierre, et al.
Publicado: (2024)
por: Boyeau, Pierre, et al.
Publicado: (2024)
IITK at SemEval-2024 Task 2: Exploring the Capabilities of LLMs for Safe Biomedical Natural Language Inference for Clinical Trials
por: Mandal, Shreyasi, et al.
Publicado: (2024)
por: Mandal, Shreyasi, et al.
Publicado: (2024)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
por: Wang, Xingyao, et al.
Publicado: (2023)
por: Wang, Xingyao, et al.
Publicado: (2023)
Reliable and Efficient Amortized Model-based Evaluation
por: Truong, Sang, et al.
Publicado: (2025)
por: Truong, Sang, et al.
Publicado: (2025)
Token-Level LLM Collaboration via FusionRoute
por: Xiong, Nuoya, et al.
Publicado: (2026)
por: Xiong, Nuoya, et al.
Publicado: (2026)
Deconfounded Causality-aware Parameter-Efficient Fine-Tuning for Problem-Solving Improvement of LLMs
por: Wang, Ruoyu, et al.
Publicado: (2024)
por: Wang, Ruoyu, et al.
Publicado: (2024)
Fine-Tuning Improves Information Conveyance in Language Models
por: Cheng, Yuwei, et al.
Publicado: (2026)
por: Cheng, Yuwei, et al.
Publicado: (2026)
Certified Robustness Under Bounded Levenshtein Distance
por: Rocamora, Elias Abad, et al.
Publicado: (2025)
por: Rocamora, Elias Abad, et al.
Publicado: (2025)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
por: Ding, Dujian, et al.
Publicado: (2024)
por: Ding, Dujian, et al.
Publicado: (2024)
KS-Lottery: Finding Certified Lottery Tickets for Multilingual Language Models
por: Yuan, Fei, et al.
Publicado: (2024)
por: Yuan, Fei, et al.
Publicado: (2024)
KcMF: A Knowledge-compliant Framework for Schema and Entity Matching with Fine-tuning-free LLMs
por: Xu, Yongqin, et al.
Publicado: (2024)
por: Xu, Yongqin, et al.
Publicado: (2024)
How to Train Data-Efficient LLMs
por: Sachdeva, Noveen, et al.
Publicado: (2024)
por: Sachdeva, Noveen, et al.
Publicado: (2024)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
por: Wang, Zige, et al.
Publicado: (2025)
por: Wang, Zige, et al.
Publicado: (2025)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
por: Yu, Xiaodong, et al.
Publicado: (2023)
por: Yu, Xiaodong, et al.
Publicado: (2023)
ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
por: Huang, Yuqing, et al.
Publicado: (2024)
por: Huang, Yuqing, et al.
Publicado: (2024)
Ejemplares similares
-
Certifying Knowledge Comprehension in LLMs
por: Chaudhary, Isha, et al.
Publicado: (2024) -
Cost-Effective Hallucination Detection for LLMs
por: Valentin, Simon, et al.
Publicado: (2024) -
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
por: Sansford, Hannah, et al.
Publicado: (2024) -
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
por: Li, Mingxuan, et al.
Publicado: (2025) -
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
por: Chung, Tsz Ting, et al.
Publicado: (2025)