CriticEval: Evaluating Large Language Model as Critic
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lan, Tian, Zhang, Wenwei, Xu, Chen, Huang, Heyan, Lin, Dahua, Chen, Kai, Mao, Xian-ling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
von: Ma, Zi-Ao, et al.
Veröffentlicht: (2025)
von: Ma, Zi-Ao, et al.
Veröffentlicht: (2025)
ANAH: Analytical Annotation of Hallucinations in Large Language Models
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024)
InternLM2.5-StepProver: Advancing Automated Theorem Proving via Critic-Guided Search
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
Training Language Models to Critique With Multi-agent Feedback
von: Lan, Tian, et al.
Veröffentlicht: (2024)
von: Lan, Tian, et al.
Veröffentlicht: (2024)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
von: Zhang, Guo-Biao, et al.
Veröffentlicht: (2026)
von: Zhang, Guo-Biao, et al.
Veröffentlicht: (2026)
TreeEval: Benchmark-Free Evaluation of Large Language Models through Tree Planning
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2024)
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2024)
DeepCritic: Deliberate Critique with Large Language Models
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
MedCalc-Eval and MedCalc-Env: Advancing Medical Calculation Capabilities of Large Language Models
von: Mao, Kangkun, et al.
Veröffentlicht: (2025)
von: Mao, Kangkun, et al.
Veröffentlicht: (2025)
CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Field
von: Bonzi, Doria, et al.
Veröffentlicht: (2025)
von: Bonzi, Doria, et al.
Veröffentlicht: (2025)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
Unveiling the Misuse Potential of Base Large Language Models via In-Context Learning
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
A Survey on Large Language Models from General Purpose to Medical Applications: Datasets, Methodologies, and Evaluations
von: Wang, Jinqiang, et al.
Veröffentlicht: (2024)
von: Wang, Jinqiang, et al.
Veröffentlicht: (2024)
RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2025)
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
von: Gu, Zhouhong, et al.
Veröffentlicht: (2024)
von: Gu, Zhouhong, et al.
Veröffentlicht: (2024)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Zhou, Xiaofeng, et al.
Veröffentlicht: (2025)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2023)
von: Prato, Gabriele, et al.
Veröffentlicht: (2023)
Beyond Metrics: A Critical Analysis of the Variability in Large Language Model Evaluation Frameworks
von: Pimentel, Marco AF, et al.
Veröffentlicht: (2024)
von: Pimentel, Marco AF, et al.
Veröffentlicht: (2024)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2024)
Enhancing Decision-Making of Large Language Models via Actor-Critic
von: Dong, Heng, et al.
Veröffentlicht: (2025)
von: Dong, Heng, et al.
Veröffentlicht: (2025)
Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models
von: Lu, Yi-Fan, et al.
Veröffentlicht: (2024)
von: Lu, Yi-Fan, et al.
Veröffentlicht: (2024)
Designing Domain-Specific Large Language Models: The Critical Role of Fine-Tuning in Public Opinion Simulation
von: Lin, Haocheng
Veröffentlicht: (2024)
von: Lin, Haocheng
Veröffentlicht: (2024)
Lost in the Source Language: How Large Language Models Evaluate the Quality of Machine Translation
von: Huang, Xu, et al.
Veröffentlicht: (2024)
von: Huang, Xu, et al.
Veröffentlicht: (2024)
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
von: Gupta, Prannaya, et al.
Veröffentlicht: (2024)
von: Gupta, Prannaya, et al.
Veröffentlicht: (2024)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
von: Sharma, Archit, et al.
Veröffentlicht: (2024)
von: Sharma, Archit, et al.
Veröffentlicht: (2024)
Large Language Model Critics for Execution-Free Evaluation of Code Changes
von: Yadavally, Aashish, et al.
Veröffentlicht: (2025)
von: Yadavally, Aashish, et al.
Veröffentlicht: (2025)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
PointLLM: Empowering Large Language Models to Understand Point Clouds
von: Xu, Runsen, et al.
Veröffentlicht: (2023)
von: Xu, Runsen, et al.
Veröffentlicht: (2023)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
LEAN-GitHub: Compiling GitHub LEAN repositories for a versatile LEAN prover
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
DND: Boosting Large Language Models with Dynamic Nested Depth
von: Chen, Tieyuan, et al.
Veröffentlicht: (2025)
von: Chen, Tieyuan, et al.
Veröffentlicht: (2025)
IndicEval: A Bilingual Indian Educational Evaluation Framework for Large Language Models
von: Bharti, Saurabh, et al.
Veröffentlicht: (2026)
von: Bharti, Saurabh, et al.
Veröffentlicht: (2026)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
Navigating the OverKill in Large Language Models
von: Shi, Chenyu, et al.
Veröffentlicht: (2024)
von: Shi, Chenyu, et al.
Veröffentlicht: (2024)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
von: Tu, Shangqing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
von: Ma, Zi-Ao, et al.
Veröffentlicht: (2025) -
ANAH: Analytical Annotation of Hallucinations in Large Language Models
von: Ji, Ziwei, et al.
Veröffentlicht: (2024) -
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
von: Gu, Yuzhe, et al.
Veröffentlicht: (2024) -
InternLM2.5-StepProver: Advancing Automated Theorem Proving via Critic-Guided Search
von: Wu, Zijian, et al.
Veröffentlicht: (2024) -
Training Language Models to Critique With Multi-agent Feedback
von: Lan, Tian, et al.
Veröffentlicht: (2024)