EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Yijie, Sun, Yuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EasyECR: A Library for Easy Implementation and Evaluation of Event Coreference Resolution Models
por: Li, Yuncong, et al.
Publicado: (2024)
por: Li, Yuncong, et al.
Publicado: (2024)
Easy Problems That LLMs Get Wrong
por: Williams, Sean, et al.
Publicado: (2024)
por: Williams, Sean, et al.
Publicado: (2024)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
por: Yehudai, Asaf, et al.
Publicado: (2025)
por: Yehudai, Asaf, et al.
Publicado: (2025)
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models
por: Ou, Yixin, et al.
Publicado: (2024)
por: Ou, Yixin, et al.
Publicado: (2024)
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
por: Wang, Chengyu, et al.
Publicado: (2025)
por: Wang, Chengyu, et al.
Publicado: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
por: Soligo, Anna, et al.
Publicado: (2026)
por: Soligo, Anna, et al.
Publicado: (2026)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
por: Zhao, Xiangyu, et al.
Publicado: (2023)
por: Zhao, Xiangyu, et al.
Publicado: (2023)
EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
por: Wang, Peng, et al.
Publicado: (2023)
por: Wang, Peng, et al.
Publicado: (2023)
Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation
por: Ledoyen, François, et al.
Publicado: (2025)
por: Ledoyen, François, et al.
Publicado: (2025)
YourBench: Easy Custom Evaluation Sets for Everyone
por: Shashidhar, Sumuk, et al.
Publicado: (2025)
por: Shashidhar, Sumuk, et al.
Publicado: (2025)
Frustratingly Easy Label Projection for Cross-lingual Transfer
por: Chen, Yang, et al.
Publicado: (2022)
por: Chen, Yang, et al.
Publicado: (2022)
Inclusive Easy-to-Read Generation for Individuals with Cognitive Impairments
por: Ledoyen, François, et al.
Publicado: (2025)
por: Ledoyen, François, et al.
Publicado: (2025)
Revisiting Generalization Across Difficulty Levels: It's Not So Easy
por: Kordi, Yeganeh, et al.
Publicado: (2025)
por: Kordi, Yeganeh, et al.
Publicado: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
por: Chan, Yik Siu, et al.
Publicado: (2025)
por: Chan, Yik Siu, et al.
Publicado: (2025)
EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models
por: Xu, Ziwen, et al.
Publicado: (2025)
por: Xu, Ziwen, et al.
Publicado: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
por: Thakur, Aman Singh, et al.
Publicado: (2024)
por: Thakur, Aman Singh, et al.
Publicado: (2024)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
por: Zhou, Weikang, et al.
Publicado: (2024)
por: Zhou, Weikang, et al.
Publicado: (2024)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
por: Hu, Jian, et al.
Publicado: (2024)
por: Hu, Jian, et al.
Publicado: (2024)
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
por: Sun, Zhiqing, et al.
Publicado: (2024)
por: Sun, Zhiqing, et al.
Publicado: (2024)
LLaVA-OneVision: Easy Visual Task Transfer
por: Li, Bo, et al.
Publicado: (2024)
por: Li, Bo, et al.
Publicado: (2024)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
por: Hase, Peter, et al.
Publicado: (2024)
por: Hase, Peter, et al.
Publicado: (2024)
EdaCSC: Two Easy Data Augmentation Methods for Chinese Spelling Correction
por: Sheng, Lei, et al.
Publicado: (2024)
por: Sheng, Lei, et al.
Publicado: (2024)
Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation
por: Yang, Jing, et al.
Publicado: (2024)
por: Yang, Jing, et al.
Publicado: (2024)
EasyRAG: Efficient Retrieval-Augmented Generation Framework for Automated Network Operations
por: Feng, Zhangchi, et al.
Publicado: (2024)
por: Feng, Zhangchi, et al.
Publicado: (2024)
EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
por: Xu, Haolei, et al.
Publicado: (2025)
por: Xu, Haolei, et al.
Publicado: (2025)
Text Adaptation to Plain Language and Easy Read via Automatic Post-Editing Cycles
por: Calleja, Jesús, et al.
Publicado: (2025)
por: Calleja, Jesús, et al.
Publicado: (2025)
Constrained Sampling for Language Models Should Be Easy: An MCMC Perspective
por: Gonzalez, Emmanuel Anaya, et al.
Publicado: (2025)
por: Gonzalez, Emmanuel Anaya, et al.
Publicado: (2025)
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
por: Parashar, Shubham, et al.
Publicado: (2025)
por: Parashar, Shubham, et al.
Publicado: (2025)
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
por: Li, Guojian, et al.
Publicado: (2025)
por: Li, Guojian, et al.
Publicado: (2025)
ESG Accountability Made Easy: DocQA at Your Service
por: Mishra, Lokesh, et al.
Publicado: (2023)
por: Mishra, Lokesh, et al.
Publicado: (2023)
Modulated Intervention Preference Optimization (MIPO): Keep the Easy, Refine the Difficult
por: Jang, Cheolhun
Publicado: (2024)
por: Jang, Cheolhun
Publicado: (2024)
Syntax Is Easy, Semantics Is Hard: Evaluating LLMs for LTL Translation
por: Danso, Priscilla Kyei, et al.
Publicado: (2026)
por: Danso, Priscilla Kyei, et al.
Publicado: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
CardiffNLP at CLEARS-2025: Prompting Large Language Models for Plain Language and Easy-to-Read Text Rewriting
por: Ayesh, Mutaz, et al.
Publicado: (2025)
por: Ayesh, Mutaz, et al.
Publicado: (2025)
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
por: Zheng, Mingqian, et al.
Publicado: (2025)
por: Zheng, Mingqian, et al.
Publicado: (2025)
Evading Data Contamination Detection for Language Models is (too) Easy
por: Dekoninck, Jasper, et al.
Publicado: (2024)
por: Dekoninck, Jasper, et al.
Publicado: (2024)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
por: Ding, Mucong, et al.
Publicado: (2024)
por: Ding, Mucong, et al.
Publicado: (2024)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
por: Liu, Zijun, et al.
Publicado: (2024)
por: Liu, Zijun, et al.
Publicado: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
por: Gambardella, Andrew, et al.
Publicado: (2024)
por: Gambardella, Andrew, et al.
Publicado: (2024)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
por: Yang, Gao, et al.
Publicado: (2025)
por: Yang, Gao, et al.
Publicado: (2025)
Ejemplares similares
-
EasyECR: A Library for Easy Implementation and Evaluation of Event Coreference Resolution Models
por: Li, Yuncong, et al.
Publicado: (2024) -
Easy Problems That LLMs Get Wrong
por: Williams, Sean, et al.
Publicado: (2024) -
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
por: Yehudai, Asaf, et al.
Publicado: (2025) -
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models
por: Ou, Yixin, et al.
Publicado: (2024) -
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
por: Wang, Chengyu, et al.
Publicado: (2025)