API Pack: A Massive Multi-Programming Language Dataset for API Call Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Zhen, Soria, Adriana Meza, Sun, Wei, Shen, Yikang, Panda, Rameswar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915151066693632
author Guo, Zhen
Soria, Adriana Meza
Sun, Wei
Shen, Yikang
Panda, Rameswar
author_facet Guo, Zhen
Soria, Adriana Meza
Sun, Wei
Shen, Yikang
Panda, Rameswar
contents We introduce API Pack, a massive multi-programming language dataset containing over one million instruction-API calls for improving the API call generation capabilities of large language models. Our evaluation highlights three key findings: First, fine-tuning on API Pack enables open-source models to outperform GPT-3.5 and GPT-4 in generating code for entirely new API calls. We show this by fine-tuning CodeLlama-13B on 20,000 Python instances from API Pack. Second, fine-tuning on a large dataset in one language, combined with smaller datasets from others, improves API generation accuracy across multiple languages. Third, we confirm the benefits of larger datasets for API generalization, as increasing fine-tuning data to one million instances enhances generalization to new APIs. To support further research, we open-source the API Pack dataset, trained model, and code at https://github.com/zguo0525/API-Pack.
format Preprint
id arxiv_https___arxiv_org_abs_2402_09615
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
Guo, Zhen
Soria, Adriana Meza
Sun, Wei
Shen, Yikang
Panda, Rameswar
Computation and Language
Artificial Intelligence
Machine Learning
We introduce API Pack, a massive multi-programming language dataset containing over one million instruction-API calls for improving the API call generation capabilities of large language models. Our evaluation highlights three key findings: First, fine-tuning on API Pack enables open-source models to outperform GPT-3.5 and GPT-4 in generating code for entirely new API calls. We show this by fine-tuning CodeLlama-13B on 20,000 Python instances from API Pack. Second, fine-tuning on a large dataset in one language, combined with smaller datasets from others, improves API generation accuracy across multiple languages. Third, we confirm the benefits of larger datasets for API generalization, as increasing fine-tuning data to one million instances enhances generalization to new APIs. To support further research, we open-source the API Pack dataset, trained model, and code at https://github.com/zguo0525/API-Pack.
title API Pack: A Massive Multi-Programming Language Dataset for API Call Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2402.09615