Guardado en:
Detalles Bibliográficos
Autores principales: Bhan, Nirav, Gupta, Shival, Manaswini, Sai, Baba, Ritik, Yadav, Narun, Desai, Hillori, Choudhary, Yash, Pawar, Aman, Shrivastava, Sarthak, Biswas, Sudipta
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:https://arxiv.org/abs/2410.17950
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912083413565440
author Bhan, Nirav
Gupta, Shival
Manaswini, Sai
Baba, Ritik
Yadav, Narun
Desai, Hillori
Choudhary, Yash
Pawar, Aman
Shrivastava, Sarthak
Biswas, Sudipta
author_facet Bhan, Nirav
Gupta, Shival
Manaswini, Sai
Baba, Ritik
Yadav, Narun
Desai, Hillori
Choudhary, Yash
Pawar, Aman
Shrivastava, Sarthak
Biswas, Sudipta
contents Large Language Models (LLMs) have shown remarkable capabilities in various domains, yet their economic impact has been limited by challenges in tool use and function calling. This paper introduces ThorV2, a novel architecture that significantly enhances LLMs' function calling abilities. We develop a comprehensive benchmark focused on HubSpot CRM operations to evaluate ThorV2 against leading models from OpenAI and Anthropic. Our results demonstrate that ThorV2 outperforms existing models in accuracy, reliability, latency, and cost efficiency for both single and multi-API calling tasks. We also show that ThorV2 is far more reliable and scales better to multistep tasks compared to traditional models. Our work offers the tantalizing possibility of more accurate function-calling compared to today's best-performing models using significantly smaller LLMs. These advancements have significant implications for the development of more capable AI assistants and the broader application of LLMs in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17950
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benchmarking Floworks against OpenAI & Anthropic: A Novel Framework for Enhanced LLM Function Calling
Bhan, Nirav
Gupta, Shival
Manaswini, Sai
Baba, Ritik
Yadav, Narun
Desai, Hillori
Choudhary, Yash
Pawar, Aman
Shrivastava, Sarthak
Biswas, Sudipta
Artificial Intelligence
Large Language Models (LLMs) have shown remarkable capabilities in various domains, yet their economic impact has been limited by challenges in tool use and function calling. This paper introduces ThorV2, a novel architecture that significantly enhances LLMs' function calling abilities. We develop a comprehensive benchmark focused on HubSpot CRM operations to evaluate ThorV2 against leading models from OpenAI and Anthropic. Our results demonstrate that ThorV2 outperforms existing models in accuracy, reliability, latency, and cost efficiency for both single and multi-API calling tasks. We also show that ThorV2 is far more reliable and scales better to multistep tasks compared to traditional models. Our work offers the tantalizing possibility of more accurate function-calling compared to today's best-performing models using significantly smaller LLMs. These advancements have significant implications for the development of more capable AI assistants and the broader application of LLMs in real-world scenarios.
title Benchmarking Floworks against OpenAI & Anthropic: A Novel Framework for Enhanced LLM Function Calling
topic Artificial Intelligence
url https://arxiv.org/abs/2410.17950