Automatic Generation of Model and Data Cards: A Step Towards Responsible AI

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Jiarui, Li, Wenkai, Jin, Zhijing, Diab, Mona
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910495001280512
author Liu, Jiarui
Li, Wenkai
Jin, Zhijing
Diab, Mona
author_facet Liu, Jiarui
Li, Wenkai
Jin, Zhijing
Diab, Mona
contents In an era of model and data proliferation in machine learning/AI especially marked by the rapid advancement of open-sourced technologies, there arises a critical need for standardized consistent documentation. Our work addresses the information incompleteness in current human-generated model and data cards. We propose an automated generation approach using Large Language Models (LLMs). Our key contributions include the establishment of CardBench, a comprehensive dataset aggregated from over 4.8k model cards and 1.4k data cards, coupled with the development of the CardGen pipeline comprising a two-step retrieval process. Our approach exhibits enhanced completeness, objectivity, and faithfulness in generated model and data cards, a significant step in responsible AI documentation practices ensuring better accountability and traceability.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06258
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Automatic Generation of Model and Data Cards: A Step Towards Responsible AI
Liu, Jiarui
Li, Wenkai
Jin, Zhijing
Diab, Mona
Computation and Language
In an era of model and data proliferation in machine learning/AI especially marked by the rapid advancement of open-sourced technologies, there arises a critical need for standardized consistent documentation. Our work addresses the information incompleteness in current human-generated model and data cards. We propose an automated generation approach using Large Language Models (LLMs). Our key contributions include the establishment of CardBench, a comprehensive dataset aggregated from over 4.8k model cards and 1.4k data cards, coupled with the development of the CardGen pipeline comprising a two-step retrieval process. Our approach exhibits enhanced completeness, objectivity, and faithfulness in generated model and data cards, a significant step in responsible AI documentation practices ensuring better accountability and traceability.
title Automatic Generation of Model and Data Cards: A Step Towards Responsible AI
topic Computation and Language
url https://arxiv.org/abs/2405.06258