Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Joshi, Ratnesh Kumar, Priya, Priyanshu, Desai, Vishesh, Dudhate, Saurav, Senapati, Siddhant, Ekbal, Asif, Ramnani, Roshni, Maitra, Anutosh, Sengupta, Shubhashis
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916497886019584
author Joshi, Ratnesh Kumar
Priya, Priyanshu
Desai, Vishesh
Dudhate, Saurav
Senapati, Siddhant
Ekbal, Asif
Ramnani, Roshni
Maitra, Anutosh
Sengupta, Shubhashis
author_facet Joshi, Ratnesh Kumar
Priya, Priyanshu
Desai, Vishesh
Dudhate, Saurav
Senapati, Siddhant
Ekbal, Asif
Ramnani, Roshni
Maitra, Anutosh
Sengupta, Shubhashis
contents Given the advancements in conversational artificial intelligence, the evaluation and assessment of Large Language Models (LLMs) play a crucial role in ensuring optimal performance across various conversational tasks. In this paper, we present a comprehensive study that thoroughly evaluates the capabilities and limitations of five prevalent LLMs: Llama, OPT, Falcon, Alpaca, and MPT. The study encompasses various conversational tasks, including reservation, empathetic response generation, mental health and legal counseling, persuasion, and negotiation. To conduct the evaluation, an extensive test setup is employed, utilizing multiple evaluation criteria that span from automatic to human evaluation. This includes using generic and task-specific metrics to gauge the LMs' performance accurately. From our evaluation, no single model emerges as universally optimal for all tasks. Instead, their performance varies significantly depending on the specific requirements of each task. While some models excel in certain tasks, they may demonstrate comparatively poorer performance in others. These findings emphasize the importance of considering task-specific requirements and characteristics when selecting the most suitable LM for conversational applications.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17204
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks
Joshi, Ratnesh Kumar
Priya, Priyanshu
Desai, Vishesh
Dudhate, Saurav
Senapati, Siddhant
Ekbal, Asif
Ramnani, Roshni
Maitra, Anutosh
Sengupta, Shubhashis
Computation and Language
Artificial Intelligence
Given the advancements in conversational artificial intelligence, the evaluation and assessment of Large Language Models (LLMs) play a crucial role in ensuring optimal performance across various conversational tasks. In this paper, we present a comprehensive study that thoroughly evaluates the capabilities and limitations of five prevalent LLMs: Llama, OPT, Falcon, Alpaca, and MPT. The study encompasses various conversational tasks, including reservation, empathetic response generation, mental health and legal counseling, persuasion, and negotiation. To conduct the evaluation, an extensive test setup is employed, utilizing multiple evaluation criteria that span from automatic to human evaluation. This includes using generic and task-specific metrics to gauge the LMs' performance accurately. From our evaluation, no single model emerges as universally optimal for all tasks. Instead, their performance varies significantly depending on the specific requirements of each task. While some models excel in certain tasks, they may demonstrate comparatively poorer performance in others. These findings emphasize the importance of considering task-specific requirements and characteristics when selecting the most suitable LM for conversational applications.
title Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.17204