Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915036238184448 |
|---|---|
| author | Fedorov, Igor Plawiak, Kate Wu, Lemeng Elgamal, Tarek Suda, Naveen Smith, Eric Zhan, Hongyuan Chi, Jianfeng Hulovatyy, Yuriy Patel, Kimish Liu, Zechun Zhao, Changsheng Shi, Yangyang Blankevoort, Tijmen Pasupuleti, Mahesh Soran, Bilge Coudert, Zacharie Delpierre Alao, Rachad Krishnamoorthi, Raghuraman Chandra, Vikas |
| author_facet | Fedorov, Igor Plawiak, Kate Wu, Lemeng Elgamal, Tarek Suda, Naveen Smith, Eric Zhan, Hongyuan Chi, Jianfeng Hulovatyy, Yuriy Patel, Kimish Liu, Zechun Zhao, Changsheng Shi, Yangyang Blankevoort, Tijmen Pasupuleti, Mahesh Soran, Bilge Coudert, Zacharie Delpierre Alao, Rachad Krishnamoorthi, Raghuraman Chandra, Vikas |
| contents | This paper presents Llama Guard 3-1B-INT4, a compact and efficient Llama Guard model, which has been open-sourced to the community during Meta Connect 2024. We demonstrate that Llama Guard 3-1B-INT4 can be deployed on resource-constrained devices, achieving a throughput of at least 30 tokens per second and a time-to-first-token of 2.5 seconds or less on a commodity Android mobile CPU. Notably, our experiments show that Llama Guard 3-1B-INT4 attains comparable or superior safety moderation scores to its larger counterpart, Llama Guard 3-1B, despite being approximately 7 times smaller in size (440MB). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_17713 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations Fedorov, Igor Plawiak, Kate Wu, Lemeng Elgamal, Tarek Suda, Naveen Smith, Eric Zhan, Hongyuan Chi, Jianfeng Hulovatyy, Yuriy Patel, Kimish Liu, Zechun Zhao, Changsheng Shi, Yangyang Blankevoort, Tijmen Pasupuleti, Mahesh Soran, Bilge Coudert, Zacharie Delpierre Alao, Rachad Krishnamoorthi, Raghuraman Chandra, Vikas Distributed, Parallel, and Cluster Computing Artificial Intelligence This paper presents Llama Guard 3-1B-INT4, a compact and efficient Llama Guard model, which has been open-sourced to the community during Meta Connect 2024. We demonstrate that Llama Guard 3-1B-INT4 can be deployed on resource-constrained devices, achieving a throughput of at least 30 tokens per second and a time-to-first-token of 2.5 seconds or less on a commodity Android mobile CPU. Notably, our experiments show that Llama Guard 3-1B-INT4 attains comparable or superior safety moderation scores to its larger counterpart, Llama Guard 3-1B, despite being approximately 7 times smaller in size (440MB). |
| title | Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations |
| topic | Distributed, Parallel, and Cluster Computing Artificial Intelligence |
| url | https://arxiv.org/abs/2411.17713 |