Fusion-Eval: Integrating Assistant Evaluators with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Shu, Lei, Wichers, Nevan, Luo, Liangchen, Zhu, Yun, Liu, Yinxiao, Chen, Jindong, Meng, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection
by: Zhu, Yun, et al.
Published: (2024)
by: Zhu, Yun, et al.
Published: (2024)
Gradient-Based Language Model Red Teaming
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024)
by: Lee, Yuho, et al.
Published: (2024)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
by: Zhao, Jiahao, et al.
Published: (2025)
by: Zhao, Jiahao, et al.
Published: (2025)
TN-Eval: Rubric and Evaluation Protocols for Measuring the Quality of Behavioral Therapy Notes
by: Shah, Raj Sanjay, et al.
Published: (2025)
by: Shah, Raj Sanjay, et al.
Published: (2025)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
by: Zhang, Mengyuan, et al.
Published: (2024)
by: Zhang, Mengyuan, et al.
Published: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?
by: Lei, Zhikai, et al.
Published: (2024)
by: Lei, Zhikai, et al.
Published: (2024)
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
by: Wang, Ganghua, et al.
Published: (2025)
by: Wang, Ganghua, et al.
Published: (2025)
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
by: Zhang, Shutong, et al.
Published: (2026)
by: Zhang, Shutong, et al.
Published: (2026)
Visualizing Neural Network Imagination
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Evaluation of Finetuned LLMs in AMR Parsing
by: Ho, Shu Han
Published: (2025)
by: Ho, Shu Han
Published: (2025)
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
by: Ma, Zi-Ao, et al.
Published: (2025)
by: Ma, Zi-Ao, et al.
Published: (2025)
CriticEval: Evaluating Large Language Model as Critic
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations
by: Li, Jiatong, et al.
Published: (2024)
by: Li, Jiatong, et al.
Published: (2024)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
by: Liu, Zhiqiang, et al.
Published: (2025)
by: Liu, Zhiqiang, et al.
Published: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
by: Jiang, Botian, et al.
Published: (2024)
by: Jiang, Botian, et al.
Published: (2024)
RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs
by: Huang, Zhongzhan, et al.
Published: (2025)
by: Huang, Zhongzhan, et al.
Published: (2025)
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
by: Wang, Yiheng, et al.
Published: (2025)
by: Wang, Yiheng, et al.
Published: (2025)
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels
by: Liu, Xiaoou, et al.
Published: (2025)
by: Liu, Xiaoou, et al.
Published: (2025)
Can an Individual Manipulate the Collective Decisions of Multi-Agents?
by: Liu, Fengyuan, et al.
Published: (2025)
by: Liu, Fengyuan, et al.
Published: (2025)
User-Assistant Bias in LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
LLMs to Support a Domain Specific Knowledge Assistant
by: Lovin, Maria-Flavia
Published: (2025)
by: Lovin, Maria-Flavia
Published: (2025)
PAD: Personalized Alignment of LLMs at Decoding-Time
by: Chen, Ruizhe, et al.
Published: (2024)
by: Chen, Ruizhe, et al.
Published: (2024)
Explore the Reasoning Capability of LLMs in the Chess Testbed
by: Wang, Shu, et al.
Published: (2024)
by: Wang, Shu, et al.
Published: (2024)
Model Spec Midtraining: Improving How Alignment Training Generalizes
by: Li, Chloe, et al.
Published: (2026)
by: Li, Chloe, et al.
Published: (2026)
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
by: Ran, Delong, et al.
Published: (2024)
by: Ran, Delong, et al.
Published: (2024)
MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024)
by: Xie, Yuxuan, et al.
Published: (2024)
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
by: Hua, Andong, et al.
Published: (2025)
by: Hua, Andong, et al.
Published: (2025)
LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning
by: Yin, Joy Lim Jia, et al.
Published: (2025)
by: Yin, Joy Lim Jia, et al.
Published: (2025)
Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
by: Balog, Krisztian, et al.
Published: (2025)
by: Balog, Krisztian, et al.
Published: (2025)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
by: Garg, Madhav Krishan, et al.
Published: (2025)
by: Garg, Madhav Krishan, et al.
Published: (2025)
Similar Items
-
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
by: Cao, Meng, et al.
Published: (2024) -
Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection
by: Zhu, Yun, et al.
Published: (2024) -
Gradient-Based Language Model Red Teaming
by: Wichers, Nevan, et al.
Published: (2024) -
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024) -
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
by: Zhao, Jiahao, et al.
Published: (2025)