BotEval: Facilitating Interactive Human Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Hyundong, Gowda, Thamme, Huang, Yuyang, Lu, Zixun, Tong, Tianli, May, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Many-to-English Machine Translation Tools, Data, and Pretrained Models
by: Gowda, Thamme, et al.
Published: (2021)
by: Gowda, Thamme, et al.
Published: (2021)
Can Language Model Moderators Improve the Health of Online Discourse?
by: Cho, Hyundong, et al.
Published: (2023)
by: Cho, Hyundong, et al.
Published: (2023)
PyMarian: Fast Neural Machine Translation and Evaluation in Python
by: Gowda, Thamme, et al.
Published: (2024)
by: Gowda, Thamme, et al.
Published: (2024)
Which Questions Improve Learning the Most? Utility Estimation of Questions with LM-based Simulations
by: Lee, Dong-Ho, et al.
Published: (2025)
by: Lee, Dong-Ho, et al.
Published: (2025)
NewsEdits 2.0: Learning the Intentions Behind Updating News
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
NewsInterview: a Dataset and a Playground to Evaluate LLMs' Ground Gap via Informational Interviews
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
Can Vision Language Models Understand Mimed Actions?
by: Cho, Hyundong, et al.
Published: (2025)
by: Cho, Hyundong, et al.
Published: (2025)
Tuning-Free Personalized Alignment via Trial-Error-Explain In-Context Learning
by: Cho, Hyundong, et al.
Published: (2025)
by: Cho, Hyundong, et al.
Published: (2025)
BatchEval: Towards Human-like Text Evaluation
by: Yuan, Peiwen, et al.
Published: (2023)
by: Yuan, Peiwen, et al.
Published: (2023)
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles
by: Huang, Shulin, et al.
Published: (2023)
by: Huang, Shulin, et al.
Published: (2023)
Continual Dialogue State Tracking via Example-Guided Question Answering
by: Cho, Hyundong, et al.
Published: (2023)
by: Cho, Hyundong, et al.
Published: (2023)
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
by: Dejl, Adam, et al.
Published: (2026)
by: Dejl, Adam, et al.
Published: (2026)
HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants
by: Gritta, Milan, et al.
Published: (2024)
by: Gritta, Milan, et al.
Published: (2024)
AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents
by: Liang, Yuanzhi, et al.
Published: (2024)
by: Liang, Yuanzhi, et al.
Published: (2024)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
by: Zhou, Lingfeng, et al.
Published: (2025)
by: Zhou, Lingfeng, et al.
Published: (2025)
AlphaEval: Evaluating Agents in Production
by: Lu, Pengrui, et al.
Published: (2026)
by: Lu, Pengrui, et al.
Published: (2026)
A Little Human Data Goes A Long Way
by: Ashok, Dhananjay, et al.
Published: (2024)
by: Ashok, Dhananjay, et al.
Published: (2024)
Sonnet or Not, Bot? Poetry Evaluation for Large Models and Datasets
by: Walsh, Melanie, et al.
Published: (2024)
by: Walsh, Melanie, et al.
Published: (2024)
mHumanEval -- A Multilingual Benchmark to Evaluate Large Language Models for Code Generation
by: Raihan, Nishat, et al.
Published: (2024)
by: Raihan, Nishat, et al.
Published: (2024)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Spot the bot: Coarse-Grained Partition of Semantic Paths for Bots and Humans
by: Gromov, Vasilii A., et al.
Published: (2024)
by: Gromov, Vasilii A., et al.
Published: (2024)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
by: Mireshghallah, Niloofar, et al.
Published: (2024)
by: Mireshghallah, Niloofar, et al.
Published: (2024)
Bot or Human? Detecting ChatGPT Imposters with A Single Question
by: Wang, Hong, et al.
Published: (2023)
by: Wang, Hong, et al.
Published: (2023)
From Human-to-Human to Human-to-Bot Conversations in Software Engineering
by: Khojah, Ranim, et al.
Published: (2024)
by: Khojah, Ranim, et al.
Published: (2024)
MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs
by: Fan, Yongqi, et al.
Published: (2025)
by: Fan, Yongqi, et al.
Published: (2025)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
by: Liu, Minqian, et al.
Published: (2023)
by: Liu, Minqian, et al.
Published: (2023)
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
by: Eisenstein, Jacob, et al.
Published: (2026)
by: Eisenstein, Jacob, et al.
Published: (2026)
Evaluating Gender Bias of LLMs in Making Morality Judgements
by: Bajaj, Divij, et al.
Published: (2024)
by: Bajaj, Divij, et al.
Published: (2024)
SocialEval: Evaluating Social Intelligence of Large Language Models
by: Zhou, Jinfeng, et al.
Published: (2025)
by: Zhou, Jinfeng, et al.
Published: (2025)
HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
by: Liu, Yuxuan, et al.
Published: (2024)
by: Liu, Yuxuan, et al.
Published: (2024)
CriticEval: Evaluating Large Language Model as Critic
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
by: Kim, Tae Soo, et al.
Published: (2023)
by: Kim, Tae Soo, et al.
Published: (2023)
Evaluating Human-Language Model Interaction
by: Lee, Mina, et al.
Published: (2022)
by: Lee, Mina, et al.
Published: (2022)
Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals
by: Akinrele, Akindoyin, et al.
Published: (2026)
by: Akinrele, Akindoyin, et al.
Published: (2026)
MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation
by: Srun, Nalin, et al.
Published: (2026)
by: Srun, Nalin, et al.
Published: (2026)
Style Transfer with Multi-iteration Preference Optimization
by: Liu, Shuai, et al.
Published: (2024)
by: Liu, Shuai, et al.
Published: (2024)
Conceptual Steganography
by: Zhou, Zhejian, et al.
Published: (2026)
by: Zhou, Zhejian, et al.
Published: (2026)
ProLex: A Benchmark for Language Proficiency-oriented Lexical Substitution
by: Zhang, Xuanming, et al.
Published: (2024)
by: Zhang, Xuanming, et al.
Published: (2024)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
by: Prato, Gabriele, et al.
Published: (2023)
by: Prato, Gabriele, et al.
Published: (2023)
Similar Items
-
Many-to-English Machine Translation Tools, Data, and Pretrained Models
by: Gowda, Thamme, et al.
Published: (2021) -
Can Language Model Moderators Improve the Health of Online Discourse?
by: Cho, Hyundong, et al.
Published: (2023) -
PyMarian: Fast Neural Machine Translation and Evaluation in Python
by: Gowda, Thamme, et al.
Published: (2024) -
Which Questions Improve Learning the Most? Utility Estimation of Questions with LM-based Simulations
by: Lee, Dong-Ho, et al.
Published: (2025) -
NewsEdits 2.0: Learning the Intentions Behind Updating News
by: Spangher, Alexander, et al.
Published: (2024)