Salvato in:
| Autori principali: | Liang, Yuanzhi, Zhu, Linchao, Yang, Yi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2401.06509 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FragRel: Exploiting Fragment-level Relations in the External Memory of Large Language Models
di: Yue, Xihang, et al.
Pubblicazione: (2024)
di: Yue, Xihang, et al.
Pubblicazione: (2024)
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
di: Lu, Yu, et al.
Pubblicazione: (2024)
di: Lu, Yu, et al.
Pubblicazione: (2024)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
di: Liang, Chumeng, et al.
Pubblicazione: (2025)
di: Liang, Chumeng, et al.
Pubblicazione: (2025)
SocialEval: Evaluating Social Intelligence of Large Language Models
di: Zhou, Jinfeng, et al.
Pubblicazione: (2025)
di: Zhou, Jinfeng, et al.
Pubblicazione: (2025)
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback
di: Chu, Seongyeub, et al.
Pubblicazione: (2026)
di: Chu, Seongyeub, et al.
Pubblicazione: (2026)
RepEval: Effective Text Evaluation with LLM Representation
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
AlphaEval: Evaluating Agents in Production
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
di: Zhao, Shuai, et al.
Pubblicazione: (2025)
di: Zhao, Shuai, et al.
Pubblicazione: (2025)
SpecEval: Evaluating Model Adherence to Behavior Specifications
di: Ahmed, Ahmed, et al.
Pubblicazione: (2025)
di: Ahmed, Ahmed, et al.
Pubblicazione: (2025)
BotEval: Facilitating Interactive Human Evaluation
di: Cho, Hyundong, et al.
Pubblicazione: (2024)
di: Cho, Hyundong, et al.
Pubblicazione: (2024)
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
di: Shen, Chengyu, et al.
Pubblicazione: (2026)
di: Shen, Chengyu, et al.
Pubblicazione: (2026)
Protecting Copyrighted Material with Unique Identifiers in Large Language Model Training
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
di: Cheng, Zhili, et al.
Pubblicazione: (2025)
di: Cheng, Zhili, et al.
Pubblicazione: (2025)
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
di: Zhu, Zhiying, et al.
Pubblicazione: (2024)
di: Zhu, Zhiying, et al.
Pubblicazione: (2024)
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
di: Wang, Jiashuo, et al.
Pubblicazione: (2025)
di: Wang, Jiashuo, et al.
Pubblicazione: (2025)
PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
di: Tang, Yixuan, et al.
Pubblicazione: (2025)
di: Tang, Yixuan, et al.
Pubblicazione: (2025)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
di: Wadhwa, Manya, et al.
Pubblicazione: (2025)
di: Wadhwa, Manya, et al.
Pubblicazione: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
AgentCollab: A Self-Evaluation-Driven Collaboration Paradigm for Efficient LLM Agents
di: Gao, Wenbo, et al.
Pubblicazione: (2026)
di: Gao, Wenbo, et al.
Pubblicazione: (2026)
Evaluating Cultural and Social Awareness of LLM Web Agents
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
di: Qiu, Haoyi, et al.
Pubblicazione: (2024)
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
di: Zhou, Yixi, et al.
Pubblicazione: (2026)
di: Zhou, Yixi, et al.
Pubblicazione: (2026)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
di: Zhou, Xuhui, et al.
Pubblicazione: (2023)
di: Zhou, Xuhui, et al.
Pubblicazione: (2023)
CreativEval: Evaluating Creativity of LLM-Based Hardware Code Generation
di: DeLorenzo, Matthew, et al.
Pubblicazione: (2024)
di: DeLorenzo, Matthew, et al.
Pubblicazione: (2024)
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
di: Wei, Tianjun, et al.
Pubblicazione: (2025)
di: Wei, Tianjun, et al.
Pubblicazione: (2025)
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
di: Dejl, Adam, et al.
Pubblicazione: (2026)
di: Dejl, Adam, et al.
Pubblicazione: (2026)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
di: Zhou, Lingfeng, et al.
Pubblicazione: (2025)
di: Zhou, Lingfeng, et al.
Pubblicazione: (2025)
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios
di: Mou, Xinyi, et al.
Pubblicazione: (2024)
di: Mou, Xinyi, et al.
Pubblicazione: (2024)
PLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games
di: Zhu, Qinglin, et al.
Pubblicazione: (2024)
di: Zhu, Qinglin, et al.
Pubblicazione: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
di: Khatun, Aisha, et al.
Pubblicazione: (2024)
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
di: Kamoi, Ryo, et al.
Pubblicazione: (2026)
di: Kamoi, Ryo, et al.
Pubblicazione: (2026)
ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding
di: Azime, Israel Abebe, et al.
Pubblicazione: (2024)
di: Azime, Israel Abebe, et al.
Pubblicazione: (2024)
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
di: Lior, Gili, et al.
Pubblicazione: (2025)
di: Lior, Gili, et al.
Pubblicazione: (2025)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
di: Zhou, Yuhang, et al.
Pubblicazione: (2025)
di: Zhou, Yuhang, et al.
Pubblicazione: (2025)
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
di: Tu, Quan, et al.
Pubblicazione: (2024)
di: Tu, Quan, et al.
Pubblicazione: (2024)
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
di: Zhang, Xiechi, et al.
Pubblicazione: (2025)
di: Zhang, Xiechi, et al.
Pubblicazione: (2025)
CiteEval: Principle-Driven Citation Evaluation for Source Attribution
di: Xu, Yumo, et al.
Pubblicazione: (2025)
di: Xu, Yumo, et al.
Pubblicazione: (2025)
MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries
di: Grolleau, François, et al.
Pubblicazione: (2025)
di: Grolleau, François, et al.
Pubblicazione: (2025)
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
di: Khan, Haidar, et al.
Pubblicazione: (2025)
di: Khan, Haidar, et al.
Pubblicazione: (2025)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
di: Zhao, Jiahao, et al.
Pubblicazione: (2025)
di: Zhao, Jiahao, et al.
Pubblicazione: (2025)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
di: Liu, Tianjian, et al.
Pubblicazione: (2025)
di: Liu, Tianjian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FragRel: Exploiting Fragment-level Relations in the External Memory of Large Language Models
di: Yue, Xihang, et al.
Pubblicazione: (2024) -
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
di: Lu, Yu, et al.
Pubblicazione: (2024) -
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
di: Liang, Chumeng, et al.
Pubblicazione: (2025) -
SocialEval: Evaluating Social Intelligence of Large Language Models
di: Zhou, Jinfeng, et al.
Pubblicazione: (2025) -
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback
di: Chu, Seongyeub, et al.
Pubblicazione: (2026)