Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition
Fuente:
arXiv
Salvato in:
| Autori principali: | Feng, Kehua, Ding, Keyan, Tan, Hongzhi, Ma, Kede, Wang, Zhihua, Guo, Shuangquan, Cheng, Yuzhou, Sun, Ge, Zheng, Guozhou, Zhang, Qiang, Chen, Huajun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions
di: Luo, Shunyang, et al.
Pubblicazione: (2026)
di: Luo, Shunyang, et al.
Pubblicazione: (2026)
Evaluating Multi-turn Human-AI Interaction
di: Ding, Shi, et al.
Pubblicazione: (2026)
di: Ding, Shi, et al.
Pubblicazione: (2026)
How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
di: Xu, Ziwen, et al.
Pubblicazione: (2026)
di: Xu, Ziwen, et al.
Pubblicazione: (2026)
Cooperative Dynamics of Censorship, Misinformation, and Influence Operations: Insights from the Global South and U.S
di: Hakami, Zaid, et al.
Pubblicazione: (2025)
di: Hakami, Zaid, et al.
Pubblicazione: (2025)
Moderating Embodied Cyber Threats Using Generative AI
di: Guo, Keyan, et al.
Pubblicazione: (2024)
di: Guo, Keyan, et al.
Pubblicazione: (2024)
Discrepancies in Mental Workload Estimation: Self-Reported versus EEG-Based Measures in Data Visualization Evaluation
di: Yim, Soobin, et al.
Pubblicazione: (2025)
di: Yim, Soobin, et al.
Pubblicazione: (2025)
Does Personalized Nudging Wear Off? A Longitudinal Study of AI Self-Modeling for Behavioral Engagement
di: He, Qing, et al.
Pubblicazione: (2026)
di: He, Qing, et al.
Pubblicazione: (2026)
Discrepancy-Aware Contrastive Adaptation in Medical Time Series Analysis
di: Wang, Yifan, et al.
Pubblicazione: (2025)
di: Wang, Yifan, et al.
Pubblicazione: (2025)
Beyond Competitive Gaming: How Casual Players Evaluate and Respond to Teammate Performance
di: Nathan, Kaushall Senthil, et al.
Pubblicazione: (2025)
di: Nathan, Kaushall Senthil, et al.
Pubblicazione: (2025)
Evaluating MEDIRL: A Replication and Ablation Study of Maximum Entropy Deep Inverse Reinforcement Learning for Human Social Navigation
di: Gupta, Vinay, et al.
Pubblicazione: (2024)
di: Gupta, Vinay, et al.
Pubblicazione: (2024)
EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models
di: Xu, Ziwen, et al.
Pubblicazione: (2025)
di: Xu, Ziwen, et al.
Pubblicazione: (2025)
AI as a Child of Mother Earth: Regrounding Human-AI Interaction in Ecological Thinking
di: Xu, Chunchen, et al.
Pubblicazione: (2024)
di: Xu, Chunchen, et al.
Pubblicazione: (2024)
Will You Participate? Exploring the Potential of Robotics Competitions on Human-centric Topics
di: Zhang, Yuchong, et al.
Pubblicazione: (2024)
di: Zhang, Yuchong, et al.
Pubblicazione: (2024)
JailbreakHunter: A Visual Analytics Approach for Jailbreak Prompts Discovery from Large-Scale Human-LLM Conversational Datasets
di: Jin, Zhihua, et al.
Pubblicazione: (2024)
di: Jin, Zhihua, et al.
Pubblicazione: (2024)
Understanding Emotional Body Expressions via Large Language Models
di: Lu, Haifeng, et al.
Pubblicazione: (2024)
di: Lu, Haifeng, et al.
Pubblicazione: (2024)
Impact of Cognitive Load on Human Trust in Hybrid Human-Robot Collaboration
di: Guo, Hao, et al.
Pubblicazione: (2024)
di: Guo, Hao, et al.
Pubblicazione: (2024)
Evaluation of Large Language Model-Driven AutoML in Data and Model Management from Human-Centered Perspective
di: Yao, Jiapeng, et al.
Pubblicazione: (2025)
di: Yao, Jiapeng, et al.
Pubblicazione: (2025)
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models
di: Ou, Yixin, et al.
Pubblicazione: (2024)
di: Ou, Yixin, et al.
Pubblicazione: (2024)
Advancing GUI for Generative AI: Charting the Design Space of Human-AI Interactions through Task Creativity and Complexity
di: Ding, Zijian
Pubblicazione: (2024)
di: Ding, Zijian
Pubblicazione: (2024)
Challenges in Trustworthy Human Evaluation of Chatbots
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
An Efficient Interaction Human-AI Synergy System Bridging Visual Awareness and Large Language Model for Intensive Care Units
di: Zhao, Yibowen, et al.
Pubblicazione: (2025)
di: Zhao, Yibowen, et al.
Pubblicazione: (2025)
FAIR: Framing AIs Role in Programming Competitions -- Understanding How LLMs Are Changing the Game in Competitive Programming
di: Pan, Dongyijie Primo, et al.
Pubblicazione: (2025)
di: Pan, Dongyijie Primo, et al.
Pubblicazione: (2025)
Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust
di: Wu, Yuzhou, et al.
Pubblicazione: (2025)
di: Wu, Yuzhou, et al.
Pubblicazione: (2025)
Development and Evaluation Study of Intelligent Cockpit in the Age of Large Models
di: Ma, Jun, et al.
Pubblicazione: (2024)
di: Ma, Jun, et al.
Pubblicazione: (2024)
Large Language Model-based Human-Agent Collaboration for Complex Task Solving
di: Feng, Xueyang, et al.
Pubblicazione: (2024)
di: Feng, Xueyang, et al.
Pubblicazione: (2024)
Large Language Model Agent Personality and Response Appropriateness: Evaluation by Human Linguistic Experts, LLM-as-Judge, and Natural Language Processing Model
di: Jayakumar, Eswari, et al.
Pubblicazione: (2025)
di: Jayakumar, Eswari, et al.
Pubblicazione: (2025)
H is for Human and How (Not) To Evaluate Qualitative Research in HCI
di: Crabtree, Andy
Pubblicazione: (2024)
di: Crabtree, Andy
Pubblicazione: (2024)
Evaluating the Influences of Explanation Style on Human-AI Reliance
di: Casolin, Emma, et al.
Pubblicazione: (2024)
di: Casolin, Emma, et al.
Pubblicazione: (2024)
Characterizing Similarities and Divergences in Conversational Tones in Humans and LLMs by Sampling with People
di: Huang, Dun-Ming, et al.
Pubblicazione: (2024)
di: Huang, Dun-Ming, et al.
Pubblicazione: (2024)
Designing Around Stigma: Human-Centered LLMs for Menstrual Health
di: Shahnawaz, Amna, et al.
Pubblicazione: (2026)
di: Shahnawaz, Amna, et al.
Pubblicazione: (2026)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
di: Si, Chenglei, et al.
Pubblicazione: (2023)
di: Si, Chenglei, et al.
Pubblicazione: (2023)
Are Humans as Brittle as Large Language Models?
di: Li, Jiahui, et al.
Pubblicazione: (2025)
di: Li, Jiahui, et al.
Pubblicazione: (2025)
Jokeasy: Exploring Human-AI Collaboration in Thematic Joke Generation
di: Ge, Yate, et al.
Pubblicazione: (2026)
di: Ge, Yate, et al.
Pubblicazione: (2026)
Typist Experiment: an Investigation of Human-to-Human Dictation via Role-play to Inform Voice-based Text Authoring
di: Liu, Can, et al.
Pubblicazione: (2024)
di: Liu, Can, et al.
Pubblicazione: (2024)
Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
di: Selitskiy, Stanislav, et al.
Pubblicazione: (2025)
di: Selitskiy, Stanislav, et al.
Pubblicazione: (2025)
Partnering with Generative AI: Experimental Evaluation of Human-Led and Model-Led Interaction in Human-AI Co-Creation
di: Maier, Sebastian, et al.
Pubblicazione: (2025)
di: Maier, Sebastian, et al.
Pubblicazione: (2025)
Strategies and Challenges of Efficient White-Box Training for Human Activity Recognition
di: Geissler, Daniel, et al.
Pubblicazione: (2024)
di: Geissler, Daniel, et al.
Pubblicazione: (2024)
Parameter-Efficient Deep Learning for Ultrasound-Based Human-Machine Interfaces
di: Lykourinas, Antonios, et al.
Pubblicazione: (2026)
di: Lykourinas, Antonios, et al.
Pubblicazione: (2026)
Visualization Generation with Large Language Models: An Evaluation
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
Gaze Archive: Enhancing Human Memory through Active Visual Logging on Smart Glasses
di: Ren, Haoxin, et al.
Pubblicazione: (2025)
di: Ren, Haoxin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating Reward Model Generalization via Pairwise Maximum Discrepancy Competitions
di: Luo, Shunyang, et al.
Pubblicazione: (2026) -
Evaluating Multi-turn Human-AI Interaction
di: Ding, Shi, et al.
Pubblicazione: (2026) -
How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
di: Xu, Ziwen, et al.
Pubblicazione: (2026) -
Cooperative Dynamics of Censorship, Misinformation, and Influence Operations: Insights from the Global South and U.S
di: Hakami, Zaid, et al.
Pubblicazione: (2025) -
Moderating Embodied Cyber Threats Using Generative AI
di: Guo, Keyan, et al.
Pubblicazione: (2024)