Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fedorov, Igor, Plawiak, Kate, Wu, Lemeng, Elgamal, Tarek, Suda, Naveen, Smith, Eric, Zhan, Hongyuan, Chi, Jianfeng, Hulovatyy, Yuriy, Patel, Kimish, Liu, Zechun, Zhao, Changsheng, Shi, Yangyang, Blankevoort, Tijmen, Pasupuleti, Mahesh, Soran, Bilge, Coudert, Zacharie Delpierre, Alao, Rachad, Krishnamoorthi, Raghuraman, Chandra, Vikas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915036238184448
author Fedorov, Igor
Plawiak, Kate
Wu, Lemeng
Elgamal, Tarek
Suda, Naveen
Smith, Eric
Zhan, Hongyuan
Chi, Jianfeng
Hulovatyy, Yuriy
Patel, Kimish
Liu, Zechun
Zhao, Changsheng
Shi, Yangyang
Blankevoort, Tijmen
Pasupuleti, Mahesh
Soran, Bilge
Coudert, Zacharie Delpierre
Alao, Rachad
Krishnamoorthi, Raghuraman
Chandra, Vikas
author_facet Fedorov, Igor
Plawiak, Kate
Wu, Lemeng
Elgamal, Tarek
Suda, Naveen
Smith, Eric
Zhan, Hongyuan
Chi, Jianfeng
Hulovatyy, Yuriy
Patel, Kimish
Liu, Zechun
Zhao, Changsheng
Shi, Yangyang
Blankevoort, Tijmen
Pasupuleti, Mahesh
Soran, Bilge
Coudert, Zacharie Delpierre
Alao, Rachad
Krishnamoorthi, Raghuraman
Chandra, Vikas
contents This paper presents Llama Guard 3-1B-INT4, a compact and efficient Llama Guard model, which has been open-sourced to the community during Meta Connect 2024. We demonstrate that Llama Guard 3-1B-INT4 can be deployed on resource-constrained devices, achieving a throughput of at least 30 tokens per second and a time-to-first-token of 2.5 seconds or less on a commodity Android mobile CPU. Notably, our experiments show that Llama Guard 3-1B-INT4 attains comparable or superior safety moderation scores to its larger counterpart, Llama Guard 3-1B, despite being approximately 7 times smaller in size (440MB).
format Preprint
id arxiv_https___arxiv_org_abs_2411_17713
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
Fedorov, Igor
Plawiak, Kate
Wu, Lemeng
Elgamal, Tarek
Suda, Naveen
Smith, Eric
Zhan, Hongyuan
Chi, Jianfeng
Hulovatyy, Yuriy
Patel, Kimish
Liu, Zechun
Zhao, Changsheng
Shi, Yangyang
Blankevoort, Tijmen
Pasupuleti, Mahesh
Soran, Bilge
Coudert, Zacharie Delpierre
Alao, Rachad
Krishnamoorthi, Raghuraman
Chandra, Vikas
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
This paper presents Llama Guard 3-1B-INT4, a compact and efficient Llama Guard model, which has been open-sourced to the community during Meta Connect 2024. We demonstrate that Llama Guard 3-1B-INT4 can be deployed on resource-constrained devices, achieving a throughput of at least 30 tokens per second and a time-to-first-token of 2.5 seconds or less on a commodity Android mobile CPU. Notably, our experiments show that Llama Guard 3-1B-INT4 attains comparable or superior safety moderation scores to its larger counterpart, Llama Guard 3-1B, despite being approximately 7 times smaller in size (440MB).
title Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2411.17713