Exploring the Latest LLMs for Leaderboard Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kabongo, Salomon, D'Souza, Jennifer, Auer, Sören
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914862386380800
author Kabongo, Salomon
D'Souza, Jennifer
Auer, Sören
author_facet Kabongo, Salomon
D'Souza, Jennifer
Auer, Sören
contents The rapid advancements in Large Language Models (LLMs) have opened new avenues for automating complex tasks in AI research. This paper investigates the efficacy of different LLMs-Mistral 7B, Llama-2, GPT-4-Turbo and GPT-4.o in extracting leaderboard information from empirical AI research articles. We explore three types of contextual inputs to the models: DocTAET (Document Title, Abstract, Experimental Setup, and Tabular Information), DocREC (Results, Experiments, and Conclusions), and DocFULL (entire document). Our comprehensive study evaluates the performance of these models in generating (Task, Dataset, Metric, Score) quadruples from research papers. The findings reveal significant insights into the strengths and limitations of each model and context type, providing valuable guidance for future AI research automation efforts.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04383
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring the Latest LLMs for Leaderboard Extraction
Kabongo, Salomon
D'Souza, Jennifer
Auer, Sören
Computation and Language
Artificial Intelligence
The rapid advancements in Large Language Models (LLMs) have opened new avenues for automating complex tasks in AI research. This paper investigates the efficacy of different LLMs-Mistral 7B, Llama-2, GPT-4-Turbo and GPT-4.o in extracting leaderboard information from empirical AI research articles. We explore three types of contextual inputs to the models: DocTAET (Document Title, Abstract, Experimental Setup, and Tabular Information), DocREC (Results, Experiments, and Conclusions), and DocFULL (entire document). Our comprehensive study evaluates the performance of these models in generating (Task, Dataset, Metric, Score) quadruples from research papers. The findings reveal significant insights into the strengths and limitations of each model and context type, providing valuable guidance for future AI research automation efforts.
title Exploring the Latest LLMs for Leaderboard Extraction
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.04383