HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations
Fuente:
arXiv
Saved in:
| Main Authors: | Sicilia, Anthony, Gates, Jennifer C., Alikhani, Malihe |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others in Conversation Forecasting
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
Accounting for Sycophancy in Language Model Uncertainty Estimation
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
by: Imai, Saki, et al.
Published: (2025)
by: Imai, Saki, et al.
Published: (2025)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
BASIL: Bayesian Assessment of Sycophancy in LLMs
by: Atwell, Katherine, et al.
Published: (2025)
by: Atwell, Katherine, et al.
Published: (2025)
Measuring How (Not Just Whether) VLMs Build Common Ground
by: Imai, Saki, et al.
Published: (2025)
by: Imai, Saki, et al.
Published: (2025)
Generating Signed Language Instructions in Large-Scale Dialogue Systems
by: İnan, Mert, et al.
Published: (2024)
by: İnan, Mert, et al.
Published: (2024)
Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI
by: Asano, Yuya, et al.
Published: (2025)
by: Asano, Yuya, et al.
Published: (2025)
An Active Learning Framework for Inclusive Generation by Large Language Models
by: Hassan, Sabit, et al.
Published: (2024)
by: Hassan, Sabit, et al.
Published: (2024)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
by: Imai, Saki, et al.
Published: (2026)
by: Imai, Saki, et al.
Published: (2026)
Active Learning for Robust and Representative LLM Generation in Safety-Critical Scenarios
by: Hassan, Sabit, et al.
Published: (2024)
by: Hassan, Sabit, et al.
Published: (2024)
Modeling Intensification for Sign Language Generation: A Computational Approach
by: İnan, Mert, et al.
Published: (2022)
by: İnan, Mert, et al.
Published: (2022)
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation
by: İnan, Mert, et al.
Published: (2025)
by: İnan, Mert, et al.
Published: (2025)
GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human
by: Lin, Yihang, et al.
Published: (2026)
by: Lin, Yihang, et al.
Published: (2026)
Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
by: Lamparth, Max, et al.
Published: (2024)
by: Lamparth, Max, et al.
Published: (2024)
Efficient Machine Translation Corpus Generation: Integrating Human-in-the-Loop Post-Editing with Large Language Models
by: Yuksel, Kamer Ali, et al.
Published: (2025)
by: Yuksel, Kamer Ali, et al.
Published: (2025)
An Interdisciplinary Approach to Human-Centered Machine Translation
by: Carpuat, Marine, et al.
Published: (2025)
by: Carpuat, Marine, et al.
Published: (2025)
Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models
by: Elangovan, Aparna, et al.
Published: (2024)
by: Elangovan, Aparna, et al.
Published: (2024)
Role-Play Zero-Shot Prompting with Large Language Models for Open-Domain Human-Machine Conversation
by: Njifenjou, Ahmed, et al.
Published: (2024)
by: Njifenjou, Ahmed, et al.
Published: (2024)
MuTSE: A Human-in-the-Loop Multi-use Text Simplification Evaluator
by: Roscan, Rares-Alexandru, et al.
Published: (2026)
by: Roscan, Rares-Alexandru, et al.
Published: (2026)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
by: Proietti, Lorenzo, et al.
Published: (2025)
by: Proietti, Lorenzo, et al.
Published: (2025)
Are Human Conversations Special? A Large Language Model Perspective
by: Jawale, Toshish, et al.
Published: (2024)
by: Jawale, Toshish, et al.
Published: (2024)
Human-centered explanation does not fit all: The interplay of sociotechnical, cognitive, and individual factors in the effect AI explanations in algorithmic decision-making
by: Ahn, Yongsu, et al.
Published: (2025)
by: Ahn, Yongsu, et al.
Published: (2025)
LLMs4SchemaDiscovery: A Human-in-the-Loop Workflow for Scientific Schema Mining with Large Language Models
by: Sadruddin, Sameer, et al.
Published: (2025)
by: Sadruddin, Sameer, et al.
Published: (2025)
Towards a Psychology of Machines: Large Language Models Predict Human Memory
by: Huff, Markus, et al.
Published: (2024)
by: Huff, Markus, et al.
Published: (2024)
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
by: Hardy, Michael
Published: (2024)
by: Hardy, Michael
Published: (2024)
HumT DumT: Measuring and controlling human-like language in LLMs
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
by: Afzal, Anum, et al.
Published: (2024)
by: Afzal, Anum, et al.
Published: (2024)
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
by: Lyu, Xinxi, et al.
Published: (2024)
by: Lyu, Xinxi, et al.
Published: (2024)
Evaluating Creative Short Story Generation in Humans and Large Language Models
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals
by: Lin, Yuxin, et al.
Published: (2025)
by: Lin, Yuxin, et al.
Published: (2025)
HARE: HumAn pRiors, a key to small language model Efficiency
by: Zhang, Lingyun, et al.
Published: (2024)
by: Zhang, Lingyun, et al.
Published: (2024)
Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility
by: Lepori, Michael A., et al.
Published: (2025)
by: Lepori, Michael A., et al.
Published: (2025)
STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking
by: Chhetri, Tek Raj, et al.
Published: (2025)
by: Chhetri, Tek Raj, et al.
Published: (2025)
Evaluating Large Language Models with Human Feedback: Establishing a Swedish Benchmark
by: Moell, Birger
Published: (2024)
by: Moell, Birger
Published: (2024)
SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
by: Daynauth, Roland, et al.
Published: (2025)
by: Daynauth, Roland, et al.
Published: (2025)
Fuzzy Fingerprinting Encoder Pre-trained Language Models for Emotion Recognition in Conversations: Human Assessment and Validity Study
by: Pereira, Patrícia, et al.
Published: (2026)
by: Pereira, Patrícia, et al.
Published: (2026)
Similar Items
-
Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others in Conversation Forecasting
by: Sicilia, Anthony, et al.
Published: (2024) -
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
by: Sicilia, Anthony, et al.
Published: (2024) -
Accounting for Sycophancy in Language Model Uncertainty Estimation
by: Sicilia, Anthony, et al.
Published: (2024) -
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
by: Imai, Saki, et al.
Published: (2025) -
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
by: Sicilia, Anthony, et al.
Published: (2024)