Archimedes-AUEB at SemEval-2024 Task 5: LLM explains Civil Procedure

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chlapanis, Odysseas S., Androutsopoulos, Ion, Galanis, Dimitrios
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929343350964224
author Chlapanis, Odysseas S.
Androutsopoulos, Ion
Galanis, Dimitrios
author_facet Chlapanis, Odysseas S.
Androutsopoulos, Ion
Galanis, Dimitrios
contents The SemEval task on Argument Reasoning in Civil Procedure is challenging in that it requires understanding legal concepts and inferring complex arguments. Currently, most Large Language Models (LLM) excelling in the legal realm are principally purposed for classification tasks, hence their reasoning rationale is subject to contention. The approach we advocate involves using a powerful teacher-LLM (ChatGPT) to extend the training dataset with explanations and generate synthetic data. The resulting data are then leveraged to fine-tune a small student-LLM. Contrary to previous work, our explanations are not directly derived from the teacher's internal knowledge. Instead they are grounded in authentic human analyses, therefore delivering a superior reasoning signal. Additionally, a new `mutation' method generates artificial data instances inspired from existing ones. We are publicly releasing the explanations as an extension to the original dataset, along with the synthetic dataset and the prompts that were used to generate both. Our system ranked 15th in the SemEval competition. It outperforms its own teacher and can produce explanations aligned with the original human analyses, as verified by legal experts.
format Preprint
id arxiv_https___arxiv_org_abs_2405_08502
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Archimedes-AUEB at SemEval-2024 Task 5: LLM explains Civil Procedure
Chlapanis, Odysseas S.
Androutsopoulos, Ion
Galanis, Dimitrios
Computation and Language
The SemEval task on Argument Reasoning in Civil Procedure is challenging in that it requires understanding legal concepts and inferring complex arguments. Currently, most Large Language Models (LLM) excelling in the legal realm are principally purposed for classification tasks, hence their reasoning rationale is subject to contention. The approach we advocate involves using a powerful teacher-LLM (ChatGPT) to extend the training dataset with explanations and generate synthetic data. The resulting data are then leveraged to fine-tune a small student-LLM. Contrary to previous work, our explanations are not directly derived from the teacher's internal knowledge. Instead they are grounded in authentic human analyses, therefore delivering a superior reasoning signal. Additionally, a new `mutation' method generates artificial data instances inspired from existing ones. We are publicly releasing the explanations as an extension to the original dataset, along with the synthetic dataset and the prompts that were used to generate both. Our system ranked 15th in the SemEval competition. It outperforms its own teacher and can produce explanations aligned with the original human analyses, as verified by legal experts.
title Archimedes-AUEB at SemEval-2024 Task 5: LLM explains Civil Procedure
topic Computation and Language
url https://arxiv.org/abs/2405.08502