HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants
Fuente:
arXiv
Saved in:
| Main Authors: | Gritta, Milan, Lampouras, Gerasimos, Iacobacci, Ignacio |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Code-Optimise: Self-Generated Preference Data for Correctness and Efficiency
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
DReSD: Dense Retrieval for Speculative Decoding
by: Gritta, Milan, et al.
Published: (2025)
by: Gritta, Milan, et al.
Published: (2025)
DRIFT: Decompose, Retrieve, Illustrate, then Formalize Theorems
by: Zhang, Meiru, et al.
Published: (2025)
by: Zhang, Meiru, et al.
Published: (2025)
Findings of the First Workshop on Simulating Conversational Intelligence in Chat
by: Graham, Yvette, et al.
Published: (2024)
by: Graham, Yvette, et al.
Published: (2024)
Mixture of Attentions For Speculative Decoding
by: Zimmer, Matthieu, et al.
Published: (2024)
by: Zimmer, Matthieu, et al.
Published: (2024)
Text-to-Code Generation with Modality-relative Pre-training
by: Christopoulou, Fenia, et al.
Published: (2024)
by: Christopoulou, Fenia, et al.
Published: (2024)
TopoAlign: A Framework for Aligning Code to Math via Topological Decomposition
by: Li, Yupei, et al.
Published: (2025)
by: Li, Yupei, et al.
Published: (2025)
Conjecturing: An Overlooked Step in Formal Mathematical Reasoning
by: Sivakumar, Jasivan Alex, et al.
Published: (2025)
by: Sivakumar, Jasivan Alex, et al.
Published: (2025)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
by: Christopoulou, Fenia, et al.
Published: (2024)
by: Christopoulou, Fenia, et al.
Published: (2024)
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs
by: Manzoor, Muhammad Arslan, et al.
Published: (2024)
by: Manzoor, Muhammad Arslan, et al.
Published: (2024)
Human-inspired Episodic Memory for Infinite Context LLMs
by: Fountas, Zafeirios, et al.
Published: (2024)
by: Fountas, Zafeirios, et al.
Published: (2024)
Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
by: Weissweiler, Leonie, et al.
Published: (2025)
by: Weissweiler, Leonie, et al.
Published: (2025)
Humans and transformer LMs: Abstraction drives language learning
by: Jian, Jasper, et al.
Published: (2026)
by: Jian, Jasper, et al.
Published: (2026)
CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs
by: Capone, Luca, et al.
Published: (2025)
by: Capone, Luca, et al.
Published: (2025)
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
by: Tudosiu, Petru-Daniel, et al.
Published: (2024)
by: Tudosiu, Petru-Daniel, et al.
Published: (2024)
Correct and Optimal: the Regular Expression Inference Challenge
by: Valizadeh, Mojtaba, et al.
Published: (2023)
by: Valizadeh, Mojtaba, et al.
Published: (2023)
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
by: Wang, Shuting, et al.
Published: (2024)
by: Wang, Shuting, et al.
Published: (2024)
BotEval: Facilitating Interactive Human Evaluation
by: Cho, Hyundong, et al.
Published: (2024)
by: Cho, Hyundong, et al.
Published: (2024)
ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models
by: Zhao, Haiquan, et al.
Published: (2024)
by: Zhao, Haiquan, et al.
Published: (2024)
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
by: Zhang, Xiechi, et al.
Published: (2025)
by: Zhang, Xiechi, et al.
Published: (2025)
BatchEval: Towards Human-like Text Evaluation
by: Yuan, Peiwen, et al.
Published: (2023)
by: Yuan, Peiwen, et al.
Published: (2023)
MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-Formula
by: Hyeon, Sieun, et al.
Published: (2024)
by: Hyeon, Sieun, et al.
Published: (2024)
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
by: Tu, Quan, et al.
Published: (2024)
by: Tu, Quan, et al.
Published: (2024)
Increasing the Thinking Budget is Not All You Need
by: Iacobacci, Ignacio, et al.
Published: (2025)
by: Iacobacci, Ignacio, et al.
Published: (2025)
NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness
by: Singhal, Manav, et al.
Published: (2024)
by: Singhal, Manav, et al.
Published: (2024)
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Addressing the Ecological Fallacy in Larger LMs with Human Context
by: Soni, Nikita, et al.
Published: (2026)
by: Soni, Nikita, et al.
Published: (2026)
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
Encoder-Decoder Framework for Interactive Free Verses with Generation with Controllable High-Quality Rhyming
by: Pasini, Tommaso, et al.
Published: (2024)
by: Pasini, Tommaso, et al.
Published: (2024)
Are BabyLMs Second Language Learners?
by: Edman, Lukas, et al.
Published: (2024)
by: Edman, Lukas, et al.
Published: (2024)
PersonaEval: Are LLM Evaluators Human Enough to Judge Role-Play?
by: Zhou, Lingfeng, et al.
Published: (2025)
by: Zhou, Lingfeng, et al.
Published: (2025)
UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities
by: Jia, Qi, et al.
Published: (2026)
by: Jia, Qi, et al.
Published: (2026)
AILS-NTUA at SemEval-2026 Task 8: Evaluating Multi-Turn RAG Conversations
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
SoftLMs: Efficient Adaptive Low-Rank Approximation of Language Models using Soft-Thresholding Mechanism
by: Bhatnagar, Priyansh, et al.
Published: (2024)
by: Bhatnagar, Priyansh, et al.
Published: (2024)
How to Make LMs Strong Node Classifiers?
by: Xu, Zhe, et al.
Published: (2024)
by: Xu, Zhe, et al.
Published: (2024)
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
by: Dubois, Yann, et al.
Published: (2024)
by: Dubois, Yann, et al.
Published: (2024)
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
by: Askari, Raha, et al.
Published: (2025)
by: Askari, Raha, et al.
Published: (2025)
A Typologically Grounded Evaluation Framework for Word Order and Morphology Sensitivity in Multilingual Masked LMs
by: Feldman, Anna, et al.
Published: (2026)
by: Feldman, Anna, et al.
Published: (2026)
A Benchmark for Deep Information Synthesis
by: Paul, Debjit, et al.
Published: (2026)
by: Paul, Debjit, et al.
Published: (2026)
Similar Items
-
Code-Optimise: Self-Generated Preference Data for Correctness and Efficiency
by: Gee, Leonidas, et al.
Published: (2024) -
DReSD: Dense Retrieval for Speculative Decoding
by: Gritta, Milan, et al.
Published: (2025) -
DRIFT: Decompose, Retrieve, Illustrate, then Formalize Theorems
by: Zhang, Meiru, et al.
Published: (2025) -
Findings of the First Workshop on Simulating Conversational Intelligence in Chat
by: Graham, Yvette, et al.
Published: (2024) -
Mixture of Attentions For Speculative Decoding
by: Zimmer, Matthieu, et al.
Published: (2024)