Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
Fuente:
arXiv
Saved in:
| Main Authors: | Fedorov, Igor, Plawiak, Kate, Wu, Lemeng, Elgamal, Tarek, Suda, Naveen, Smith, Eric, Zhan, Hongyuan, Chi, Jianfeng, Hulovatyy, Yuriy, Patel, Kimish, Liu, Zechun, Zhao, Changsheng, Shi, Yangyang, Blankevoort, Tijmen, Pasupuleti, Mahesh, Soran, Bilge, Coudert, Zacharie Delpierre, Alao, Rachad, Krishnamoorthi, Raghuraman, Chandra, Vikas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
by: Chi, Jianfeng, et al.
Published: (2024)
by: Chi, Jianfeng, et al.
Published: (2024)
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
SqueezeSAM: User friendly mobile interactive segmentation
by: Varadarajan, Balakrishnan, et al.
Published: (2023)
by: Varadarajan, Balakrishnan, et al.
Published: (2023)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025)
by: Liu, Zechun, et al.
Published: (2025)
Efficient Track Anything
by: Xiong, Yunyang, et al.
Published: (2024)
by: Xiong, Yunyang, et al.
Published: (2024)
MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment
by: Huang, Hanxian, et al.
Published: (2026)
by: Huang, Hanxian, et al.
Published: (2026)
EdgeTAM: On-Device Track Anything Model
by: Zhou, Chong, et al.
Published: (2025)
by: Zhou, Chong, et al.
Published: (2025)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
by: Shen, Xiaoqian, et al.
Published: (2024)
by: Shen, Xiaoqian, et al.
Published: (2024)
PathFusion: Path-consistent Lidar-Camera Deep Feature Fusion
by: Wu, Lemeng, et al.
Published: (2022)
by: Wu, Lemeng, et al.
Published: (2022)
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
by: Kopiczko, Dawid J., et al.
Published: (2024)
by: Kopiczko, Dawid J., et al.
Published: (2024)
VeRA: Vector-based Random Matrix Adaptation
by: Kopiczko, Dawid J., et al.
Published: (2023)
by: Kopiczko, Dawid J., et al.
Published: (2023)
Communication Efficient Distributed Training with Distributed Lion
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
MobileMoE: Scaling On-Device Mixture of Experts
by: Chen, Yanbei, et al.
Published: (2026)
by: Chen, Yanbei, et al.
Published: (2026)
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
by: Kopiczko, Dawid J., et al.
Published: (2026)
by: Kopiczko, Dawid J., et al.
Published: (2026)
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
dTRPO: Trajectory Reduction in Policy Optimization of Diffusion Large Language Models
by: Zhang, Wenxuan, et al.
Published: (2026)
by: Zhang, Wenxuan, et al.
Published: (2026)
Pruning vs Quantization: Which is Better?
by: Kuzmin, Andrey, et al.
Published: (2023)
by: Kuzmin, Andrey, et al.
Published: (2023)
WinQ: Accelerating Quantization-Aware Training of Language Models Around Saddle Points
by: Li, Dongyue, et al.
Published: (2026)
by: Li, Dongyue, et al.
Published: (2026)
Small Vision-Language Models are Smart Compressors for Long Video Understanding
by: Fei, Junjie, et al.
Published: (2026)
by: Fei, Junjie, et al.
Published: (2026)
MobileLLM-Pro Technical Report
by: Huber, Patrick, et al.
Published: (2025)
by: Huber, Patrick, et al.
Published: (2025)
FP8 Quantization: The Power of the Exponent
by: Kuzmin, Andrey, et al.
Published: (2022)
by: Kuzmin, Andrey, et al.
Published: (2022)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
Elastic ViTs from Pretrained Models without Retraining
by: Simoncini, Walter, et al.
Published: (2025)
by: Simoncini, Walter, et al.
Published: (2025)
InterroGate: Learning to Share, Specialize, and Prune Representations for Multi-task Learning
by: Bejnordi, Babak Ehteshami, et al.
Published: (2024)
by: Bejnordi, Babak Ehteshami, et al.
Published: (2024)
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
by: Zhao, Changsheng, et al.
Published: (2025)
by: Zhao, Changsheng, et al.
Published: (2025)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
by: Chen, Shimao, et al.
Published: (2024)
by: Chen, Shimao, et al.
Published: (2024)
Physics-Informed Residual Learning for Safe and Adaptive Battery Charging Under Extreme Conditions
by: Pasupuleti, Ramakrishna
Published: (2026)
by: Pasupuleti, Ramakrishna
Published: (2026)
A K–R Constant–Based Framework for Predictive Stabilization, Uncertainty Regulation, and Physical Reservoir Computing
by: Pasupuleti, Ramakrishna
Published: (2026)
by: Pasupuleti, Ramakrishna
Published: (2026)
KR-Regulated Nonlinear Parabolic and Fractional Evolution Equations: Attractor Scaling Laws, Critical Thresholds, and Computational Efficiency
by: Pasupuleti, Ramakrishna
Published: (2026)
by: Pasupuleti, Ramakrishna
Published: (2026)
Pre Seismic Quiescence and Dynamical Regime Transitions in the Japan and Chile Earthquake Catalogs Evidence from KR Critical Slowing Down Indicators
by: Pasupuleti, Ramakrishna
Published: (2026)
by: Pasupuleti, Ramakrishna
Published: (2026)
Dynamical Solution to the Eta Problem in Spectator Field Models
by: Elgamal, Sana, et al.
Published: (2025)
by: Elgamal, Sana, et al.
Published: (2025)
Esboço para um método projetual para a complexidade do design contemporâneo
by: Rui S. D. Alão
Published: (2022)
by: Rui S. D. Alão
Published: (2022)
Sobre a complexidade dos problemas contemporâneos de design
by: Rui S. D. Alão
Published: (2020)
by: Rui S. D. Alão
Published: (2020)
The LLM Surgeon
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2023)
Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
by: Yu, Zhuoran, et al.
Published: (2023)
by: Yu, Zhuoran, et al.
Published: (2023)
Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation
by: Gao, Yilan, et al.
Published: (2026)
by: Gao, Yilan, et al.
Published: (2026)
Agent-as-a-Judge: Evaluate Agents with Agents
by: Zhuge, Mingchen, et al.
Published: (2024)
by: Zhuge, Mingchen, et al.
Published: (2024)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
by: Chehbouni, Khaoula, et al.
Published: (2024)
by: Chehbouni, Khaoula, et al.
Published: (2024)
A Unified Two-Parameter K–R Framework for Generalized Jensen, Hölder, Young, and Stability Inequalities
by: Pasupuleti, RamaKrishna
Published: (2026)
by: Pasupuleti, RamaKrishna
Published: (2026)
FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 Precision
by: Ma, Jingxiao, et al.
Published: (2025)
by: Ma, Jingxiao, et al.
Published: (2025)
Similar Items
-
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
by: Chi, Jianfeng, et al.
Published: (2024) -
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024) -
SqueezeSAM: User friendly mobile interactive segmentation
by: Varadarajan, Balakrishnan, et al.
Published: (2023) -
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025) -
Efficient Track Anything
by: Xiong, Yunyang, et al.
Published: (2024)