Granite Guardian
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909430756409344 |
|---|---|
| author | Padhi, Inkit Nagireddy, Manish Cornacchia, Giandomenico Chaudhury, Subhajit Pedapati, Tejaswini Dognin, Pierre Murugesan, Keerthiram Miehling, Erik Cooper, Martín Santillán Fraser, Kieran Zizzo, Giulio Hameed, Muhammad Zaid Purcell, Mark Desmond, Michael Pan, Qian Ashktorab, Zahra Vejsbjerg, Inge Daly, Elizabeth M. Hind, Michael Geyer, Werner Rawat, Ambrish Varshney, Kush R. Sattigeri, Prasanna |
| author_facet | Padhi, Inkit Nagireddy, Manish Cornacchia, Giandomenico Chaudhury, Subhajit Pedapati, Tejaswini Dognin, Pierre Murugesan, Keerthiram Miehling, Erik Cooper, Martín Santillán Fraser, Kieran Zizzo, Giulio Hameed, Muhammad Zaid Purcell, Mark Desmond, Michael Pan, Qian Ashktorab, Zahra Vejsbjerg, Inge Daly, Elizabeth M. Hind, Michael Geyer, Werner Rawat, Ambrish Varshney, Kush R. Sattigeri, Prasanna |
| contents | We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with any large language model (LLM). These models offer comprehensive coverage across multiple risk dimensions, including social bias, profanity, violence, sexual content, unethical behavior, jailbreaking, and hallucination-related risks such as context relevance, groundedness, and answer relevance for retrieval-augmented generation (RAG). Trained on a unique dataset combining human annotations from diverse sources and synthetic data, Granite Guardian models address risks typically overlooked by traditional risk detection models, such as jailbreaks and RAG-specific issues. With AUC scores of 0.871 and 0.854 on harmful content and RAG-hallucination-related benchmarks respectively, Granite Guardian is the most generalizable and competitive model available in the space. Released as open-source, Granite Guardian aims to promote responsible AI development across the community.
https://github.com/ibm-granite/granite-guardian |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_07724 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Granite Guardian Padhi, Inkit Nagireddy, Manish Cornacchia, Giandomenico Chaudhury, Subhajit Pedapati, Tejaswini Dognin, Pierre Murugesan, Keerthiram Miehling, Erik Cooper, Martín Santillán Fraser, Kieran Zizzo, Giulio Hameed, Muhammad Zaid Purcell, Mark Desmond, Michael Pan, Qian Ashktorab, Zahra Vejsbjerg, Inge Daly, Elizabeth M. Hind, Michael Geyer, Werner Rawat, Ambrish Varshney, Kush R. Sattigeri, Prasanna Computation and Language We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with any large language model (LLM). These models offer comprehensive coverage across multiple risk dimensions, including social bias, profanity, violence, sexual content, unethical behavior, jailbreaking, and hallucination-related risks such as context relevance, groundedness, and answer relevance for retrieval-augmented generation (RAG). Trained on a unique dataset combining human annotations from diverse sources and synthetic data, Granite Guardian models address risks typically overlooked by traditional risk detection models, such as jailbreaks and RAG-specific issues. With AUC scores of 0.871 and 0.854 on harmful content and RAG-hallucination-related benchmarks respectively, Granite Guardian is the most generalizable and competitive model available in the space. Released as open-source, Granite Guardian aims to promote responsible AI development across the community. https://github.com/ibm-granite/granite-guardian |
| title | Granite Guardian |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2412.07724 |