AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Minbeom, Lee, Hwanhee, Park, Joonsuk, Lee, Hwaran, Jung, Kyomin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909471582715904
author Kim, Minbeom
Lee, Hwanhee
Park, Joonsuk
Lee, Hwaran
Jung, Kyomin
author_facet Kim, Minbeom
Lee, Hwanhee
Park, Joonsuk
Lee, Hwaran
Jung, Kyomin
contents As the integration of large language models into daily life is on the rise, there is a clear gap in benchmarks for advising on subjective and personal dilemmas. To address this, we introduce AdvisorQA, the first benchmark developed to assess LLMs' capability in offering advice for deeply personalized concerns, utilizing the LifeProTips subreddit forum. This forum features a dynamic interaction where users post advice-seeking questions, receiving an average of 8.9 advice per query, with 164.2 upvotes from hundreds of users, embodying a collective intelligence framework. Therefore, we've completed a benchmark encompassing daily life questions, diverse corresponding responses, and majority vote ranking to train our helpfulness metric. Baseline experiments validate the efficacy of AdvisorQA through our helpfulness metric, GPT-4, and human evaluation, analyzing phenomena beyond the trade-off between helpfulness and harmlessness. AdvisorQA marks a significant leap in enhancing QA systems for providing personalized, empathetic advice, showcasing LLMs' improved understanding of human subjectivity.
format Preprint
id arxiv_https___arxiv_org_abs_2404_11826
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence
Kim, Minbeom
Lee, Hwanhee
Park, Joonsuk
Lee, Hwaran
Jung, Kyomin
Computation and Language
As the integration of large language models into daily life is on the rise, there is a clear gap in benchmarks for advising on subjective and personal dilemmas. To address this, we introduce AdvisorQA, the first benchmark developed to assess LLMs' capability in offering advice for deeply personalized concerns, utilizing the LifeProTips subreddit forum. This forum features a dynamic interaction where users post advice-seeking questions, receiving an average of 8.9 advice per query, with 164.2 upvotes from hundreds of users, embodying a collective intelligence framework. Therefore, we've completed a benchmark encompassing daily life questions, diverse corresponding responses, and majority vote ranking to train our helpfulness metric. Baseline experiments validate the efficacy of AdvisorQA through our helpfulness metric, GPT-4, and human evaluation, analyzing phenomena beyond the trade-off between helpfulness and harmlessness. AdvisorQA marks a significant leap in enhancing QA systems for providing personalized, empathetic advice, showcasing LLMs' improved understanding of human subjectivity.
title AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence
topic Computation and Language
url https://arxiv.org/abs/2404.11826