Using Large Language Models to Generate, Validate, and Apply User Intent Taxonomies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shah, Chirag, White, Ryen W., Andersen, Reid, Buscher, Georg, Counts, Scott, Das, Sarkar Snigdha Sarathi, Montazer, Ali, Manivannan, Sathish, Neville, Jennifer, Ni, Xiaochuan, Rangan, Nagu, Safavi, Tara, Suri, Siddharth, Wan, Mengting, Wang, Leijie, Yang, Longqi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911871951437824
author Shah, Chirag
White, Ryen W.
Andersen, Reid
Buscher, Georg
Counts, Scott
Das, Sarkar Snigdha Sarathi
Montazer, Ali
Manivannan, Sathish
Neville, Jennifer
Ni, Xiaochuan
Rangan, Nagu
Safavi, Tara
Suri, Siddharth
Wan, Mengting
Wang, Leijie
Yang, Longqi
author_facet Shah, Chirag
White, Ryen W.
Andersen, Reid
Buscher, Georg
Counts, Scott
Das, Sarkar Snigdha Sarathi
Montazer, Ali
Manivannan, Sathish
Neville, Jennifer
Ni, Xiaochuan
Rangan, Nagu
Safavi, Tara
Suri, Siddharth
Wan, Mengting
Wang, Leijie
Yang, Longqi
contents Log data can reveal valuable information about how users interact with Web search services, what they want, and how satisfied they are. However, analyzing user intents in log data is not easy, especially for emerging forms of Web search such as AI-driven chat. To understand user intents from log data, we need a way to label them with meaningful categories that capture their diversity and dynamics. Existing methods rely on manual or machine-learned labeling, which are either expensive or inflexible for large and dynamic datasets. We propose a novel solution using large language models (LLMs), which can generate rich and relevant concepts, descriptions, and examples for user intents. However, using LLMs to generate a user intent taxonomy and apply it for log analysis can be problematic for two main reasons: (1) such a taxonomy is not externally validated; and (2) there may be an undesirable feedback loop. To address this, we propose a new methodology with human experts and assessors to verify the quality of the LLM-generated taxonomy. We also present an end-to-end pipeline that uses an LLM with human-in-the-loop to produce, refine, and apply labels for user intent analysis in log data. We demonstrate its effectiveness by uncovering new insights into user intents from search and chat logs from the Microsoft Bing commercial search engine. The proposed work's novelty stems from the method for generating purpose-driven user intent taxonomies with strong validation. This method not only helps remove methodological and practical bottlenecks from intent-focused research, but also provides a new framework for generating, validating, and applying other kinds of taxonomies in a scalable and adaptable way with reasonable human effort.
format Preprint
id arxiv_https___arxiv_org_abs_2309_13063
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Using Large Language Models to Generate, Validate, and Apply User Intent Taxonomies
Shah, Chirag
White, Ryen W.
Andersen, Reid
Buscher, Georg
Counts, Scott
Das, Sarkar Snigdha Sarathi
Montazer, Ali
Manivannan, Sathish
Neville, Jennifer
Ni, Xiaochuan
Rangan, Nagu
Safavi, Tara
Suri, Siddharth
Wan, Mengting
Wang, Leijie
Yang, Longqi
Information Retrieval
Artificial Intelligence
Computation and Language
Log data can reveal valuable information about how users interact with Web search services, what they want, and how satisfied they are. However, analyzing user intents in log data is not easy, especially for emerging forms of Web search such as AI-driven chat. To understand user intents from log data, we need a way to label them with meaningful categories that capture their diversity and dynamics. Existing methods rely on manual or machine-learned labeling, which are either expensive or inflexible for large and dynamic datasets. We propose a novel solution using large language models (LLMs), which can generate rich and relevant concepts, descriptions, and examples for user intents. However, using LLMs to generate a user intent taxonomy and apply it for log analysis can be problematic for two main reasons: (1) such a taxonomy is not externally validated; and (2) there may be an undesirable feedback loop. To address this, we propose a new methodology with human experts and assessors to verify the quality of the LLM-generated taxonomy. We also present an end-to-end pipeline that uses an LLM with human-in-the-loop to produce, refine, and apply labels for user intent analysis in log data. We demonstrate its effectiveness by uncovering new insights into user intents from search and chat logs from the Microsoft Bing commercial search engine. The proposed work's novelty stems from the method for generating purpose-driven user intent taxonomies with strong validation. This method not only helps remove methodological and practical bottlenecks from intent-focused research, but also provides a new framework for generating, validating, and applying other kinds of taxonomies in a scalable and adaptable way with reasonable human effort.
title Using Large Language Models to Generate, Validate, and Apply User Intent Taxonomies
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2309.13063