Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | Bianchi, Federico, Suzgun, Mirac, Attanasio, Giuseppe, Röttger, Paul, Jurafsky, Dan, Hashimoto, Tatsunori, Zou, James |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
by: Suzgun, Mirac, et al.
Published: (2025)
by: Suzgun, Mirac, et al.
Published: (2025)
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
by: Di Palma, Dario, et al.
Published: (2025)
by: Di Palma, Dario, et al.
Published: (2025)
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)
Evaluating Commercial AI Chatbots as News Intermediaries
by: Suzgun, Mirac, et al.
Published: (2026)
by: Suzgun, Mirac, et al.
Published: (2026)
A Benchmark for Learning to Translate a New Language from One Grammar Book
by: Tanzer, Garrett, et al.
Published: (2023)
by: Tanzer, Garrett, et al.
Published: (2023)
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
by: Xing, Bohao, et al.
Published: (2024)
by: Xing, Bohao, et al.
Published: (2024)
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
by: Yang, Meng, et al.
Published: (2026)
by: Yang, Meng, et al.
Published: (2026)
LLaMA-E: Empowering E-commerce Authoring with Object-Interleaved Instruction Following
by: Shi, Kaize, et al.
Published: (2023)
by: Shi, Kaize, et al.
Published: (2023)
Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
by: Dahl, Matthew, et al.
Published: (2024)
by: Dahl, Matthew, et al.
Published: (2024)
LLaMA Pro: Progressive LLaMA with Block Expansion
by: Wu, Chengyue, et al.
Published: (2024)
by: Wu, Chengyue, et al.
Published: (2024)
Cost-of-Pass: An Economic Framework for Evaluating Language Models
by: Erol, Mehmet Hamza, et al.
Published: (2025)
by: Erol, Mehmet Hamza, et al.
Published: (2025)
LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning
by: Jahangir, Md. Zihad Bin, et al.
Published: (2025)
by: Jahangir, Md. Zihad Bin, et al.
Published: (2025)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
by: Dialameh, Maryam, et al.
Published: (2025)
by: Dialameh, Maryam, et al.
Published: (2025)
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
by: Kwon, Yongchan, et al.
Published: (2025)
by: Kwon, Yongchan, et al.
Published: (2025)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
by: Pernisi, Fabio, et al.
Published: (2024)
by: Pernisi, Fabio, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
by: Gema, Aryo Pradipta, et al.
Published: (2023)
by: Gema, Aryo Pradipta, et al.
Published: (2023)
LLaMA-Reg: Using LLaMA 2 for Unsupervised Medical Image Registration
by: Ma, Mingrui, et al.
Published: (2024)
by: Ma, Mingrui, et al.
Published: (2024)
Locality Alignment Improves Vision-Language Models
by: Covert, Ian, et al.
Published: (2024)
by: Covert, Ian, et al.
Published: (2024)
A Bitter Lesson for Data Filtering
by: Mohri, Christopher, et al.
Published: (2026)
by: Mohri, Christopher, et al.
Published: (2026)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
by: Zhu, Tong, et al.
Published: (2024)
by: Zhu, Tong, et al.
Published: (2024)
Do Language Models Know When They're Hallucinating References?
by: Agrawal, Ayush, et al.
Published: (2023)
by: Agrawal, Ayush, et al.
Published: (2023)
Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
by: Aftab, Danyal, et al.
Published: (2024)
by: Aftab, Danyal, et al.
Published: (2024)
How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
by: Bianchi, Federico, et al.
Published: (2024)
by: Bianchi, Federico, et al.
Published: (2024)
Adapting LLaMA Decoder to Vision Transformer
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
BanglaLlama: LLaMA for Bangla Language
by: Zehady, Abdullah Khan, et al.
Published: (2024)
by: Zehady, Abdullah Khan, et al.
Published: (2024)
Me LLaMA: Foundation Large Language Models for Medical Applications
by: Xie, Qianqian, et al.
Published: (2024)
by: Xie, Qianqian, et al.
Published: (2024)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
by: Fang, Qingkai, et al.
Published: (2024)
by: Fang, Qingkai, et al.
Published: (2024)
CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
by: Bartelds, Martijn, et al.
Published: (2025)
by: Bartelds, Martijn, et al.
Published: (2025)
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
by: Qu, Xiaoye, et al.
Published: (2024)
by: Qu, Xiaoye, et al.
Published: (2024)
Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset
by: Goldzycher, Janis, et al.
Published: (2024)
by: Goldzycher, Janis, et al.
Published: (2024)
LLaVA-c: Continual Improved Visual Instruction Tuning
by: Liu, Wenzhuo, et al.
Published: (2025)
by: Liu, Wenzhuo, et al.
Published: (2025)
LogLLaMA: Transformer-based log anomaly detection with LLaMA
by: Yang, Zhuoyi, et al.
Published: (2025)
by: Yang, Zhuoyi, et al.
Published: (2025)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
by: Chu, Xiangxiang, et al.
Published: (2024)
by: Chu, Xiangxiang, et al.
Published: (2024)
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages
by: Andersland, Michael
Published: (2024)
by: Andersland, Michael
Published: (2024)
Similar Items
-
Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
by: Suzgun, Mirac, et al.
Published: (2025) -
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
by: Suzgun, Mirac, et al.
Published: (2024) -
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023) -
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
by: Di Palma, Dario, et al.
Published: (2025) -
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
by: Plaza-del-Arco, Flor Miriam, et al.
Published: (2024)