Saved in:
Bibliographic Details
Main Authors: Bora, Adriana Eufrosina, St-Charles, Pierre-Luc, Bronzi, Mirko, Tchango, Arsène Fansi, Rousseau, Bruno, Mengersen, Kerrie
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.07022
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910894354595840
author Bora, Adriana Eufrosina
St-Charles, Pierre-Luc
Bronzi, Mirko
Tchango, Arsène Fansi
Rousseau, Bruno
Mengersen, Kerrie
author_facet Bora, Adriana Eufrosina
St-Charles, Pierre-Luc
Bronzi, Mirko
Tchango, Arsène Fansi
Rousseau, Bruno
Mengersen, Kerrie
contents Despite over a decade of legislative efforts to address modern slavery in the supply chains of large corporations, the effectiveness of government oversight remains hampered by the challenge of scrutinizing thousands of statements annually. While Large Language Models (LLMs) can be considered a well established solution for the automatic analysis and summarization of documents, recognizing concrete modern slavery countermeasures taken by companies and differentiating those from vague claims remains a challenging task. To help evaluate and fine-tune LLMs for the assessment of corporate statements, we introduce a dataset composed of 5,731 modern slavery statements taken from the Australian Modern Slavery Register and annotated at the sentence level. This paper details the construction steps for the dataset that include the careful design of annotation specifications, the selection and preprocessing of statements, and the creation of high-quality annotation subsets for effective model evaluations. To demonstrate our dataset's utility, we propose a machine learning methodology for the detection of sentences relevant to mandatory reporting requirements set by the Australian Modern Slavery Act. We then follow this methodology to benchmark modern language models under zero-shot and supervised learning settings.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07022
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements
Bora, Adriana Eufrosina
St-Charles, Pierre-Luc
Bronzi, Mirko
Tchango, Arsène Fansi
Rousseau, Bruno
Mengersen, Kerrie
Computation and Language
Artificial Intelligence
Machine Learning
Despite over a decade of legislative efforts to address modern slavery in the supply chains of large corporations, the effectiveness of government oversight remains hampered by the challenge of scrutinizing thousands of statements annually. While Large Language Models (LLMs) can be considered a well established solution for the automatic analysis and summarization of documents, recognizing concrete modern slavery countermeasures taken by companies and differentiating those from vague claims remains a challenging task. To help evaluate and fine-tune LLMs for the assessment of corporate statements, we introduce a dataset composed of 5,731 modern slavery statements taken from the Australian Modern Slavery Register and annotated at the sentence level. This paper details the construction steps for the dataset that include the careful design of annotation specifications, the selection and preprocessing of statements, and the creation of high-quality annotation subsets for effective model evaluations. To demonstrate our dataset's utility, we propose a machine learning methodology for the detection of sentences relevant to mandatory reporting requirements set by the Australian Modern Slavery Act. We then follow this methodology to benchmark modern language models under zero-shot and supervised learning settings.
title AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.07022