ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Qihuang, Ding, Liang, Liu, Juhua, Du, Bo, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PANDA: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation
by: Zhong, Qihuang, et al.
Published: (2022)
by: Zhong, Qihuang, et al.
Published: (2022)
E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation
by: Zhong, Qihuang, et al.
Published: (2022)
by: Zhong, Qihuang, et al.
Published: (2022)
Revisiting Knowledge Distillation for Autoregressive Language Models
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
KaFT: Knowledge-aware Fine-tuning for Boosting LLMs' Domain-specific Question-Answering Performance
by: Zhong, Qihuang, et al.
Published: (2025)
by: Zhong, Qihuang, et al.
Published: (2025)
Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
by: Zhong, Qihuang, et al.
Published: (2026)
by: Zhong, Qihuang, et al.
Published: (2026)
Resolving Knowledge Conflicts in Domain-specific Data Selection: A Case Study on Medical Instruction-tuning
by: Zhong, Qihuang, et al.
Published: (2025)
by: Zhong, Qihuang, et al.
Published: (2025)
Learning from Imperfect Data: Towards Efficient Knowledge Distillation of Autoregressive Language Models for Text-to-SQL
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
Try, Check and Retry: A Divide-and-Conquer Framework for Boosting Long-context Tool-Calling Performance of LLMs
by: Chen, Kunfeng, et al.
Published: (2026)
by: Chen, Kunfeng, et al.
Published: (2026)
Iterative Data Generation with Large Language Models for Aspect-based Sentiment Analysis
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
by: Ye, Maoyuan, et al.
Published: (2025)
by: Ye, Maoyuan, et al.
Published: (2025)
Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
by: Zhong, Qihuang, et al.
Published: (2026)
by: Zhong, Qihuang, et al.
Published: (2026)
ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering
by: Zhu, Yikai, et al.
Published: (2026)
by: Zhu, Yikai, et al.
Published: (2026)
Revisiting Catastrophic Forgetting in Large Language Model Tuning
by: Li, Hongyu, et al.
Published: (2024)
by: Li, Hongyu, et al.
Published: (2024)
SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
by: Lin, Jiacheng, et al.
Published: (2025)
by: Lin, Jiacheng, et al.
Published: (2025)
Consciousness Doesn't Do That
by: Matthias Michel
Published: (2026)
by: Matthias Michel
Published: (2026)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
by: Wang, Xintong, et al.
Published: (2024)
by: Wang, Xintong, et al.
Published: (2024)
OOP: Object-Oriented Programming Evaluation Benchmark for Large Language Models
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Building Accurate Translation-Tailored LLMs with Language Aware Instruction Tuning
by: Zan, Changtong, et al.
Published: (2024)
by: Zan, Changtong, et al.
Published: (2024)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
by: He, Qianxi, et al.
Published: (2025)
by: He, Qianxi, et al.
Published: (2025)
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
by: Lu, Qingyu, et al.
Published: (2023)
by: Lu, Qingyu, et al.
Published: (2023)
Enhancing Input-Label Mapping in In-Context Learning with Contrastive Decoding
by: Peng, Keqin, et al.
Published: (2025)
by: Peng, Keqin, et al.
Published: (2025)
Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't
by: Taguchi, Chihiro, et al.
Published: (2024)
by: Taguchi, Chihiro, et al.
Published: (2024)
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
by: Berlot-Attwell, Ian, et al.
Published: (2024)
by: Berlot-Attwell, Ian, et al.
Published: (2024)
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
by: Li, Victoria R., et al.
Published: (2024)
by: Li, Victoria R., et al.
Published: (2024)
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
by: Dang, Quy-Anh, et al.
Published: (2025)
by: Dang, Quy-Anh, et al.
Published: (2025)
Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
by: Gong, Shuzhi, et al.
Published: (2026)
by: Gong, Shuzhi, et al.
Published: (2026)
Improving Complex Reasoning over Knowledge Graph with Logic-Aware Curriculum Tuning
by: Xia, Tianle, et al.
Published: (2024)
by: Xia, Tianle, et al.
Published: (2024)
Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning
by: Huang, Wenke, et al.
Published: (2024)
by: Huang, Wenke, et al.
Published: (2024)
Synthesizing Instruction-Tuning Datasets with Contrastive Decoding
by: Ichinose, Tatsuya, et al.
Published: (2026)
by: Ichinose, Tatsuya, et al.
Published: (2026)
Just Because We Camp, Doesn't Mean We Should: The Ethics of Modelling Queer Voices
by: Sigurgeirsson, Atli, et al.
Published: (2024)
by: Sigurgeirsson, Atli, et al.
Published: (2024)
Teach AI What It Doesn't Know
by: Sean Du
Published: (2026)
by: Sean Du
Published: (2026)
IAPT: Instruction-Aware Prompt Tuning for Large Language Models
by: Zhu, Wei, et al.
Published: (2024)
by: Zhu, Wei, et al.
Published: (2024)
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language
by: Chance, Christina, et al.
Published: (2026)
by: Chance, Christina, et al.
Published: (2026)
Investigating Instruction Tuning Large Language Models on Graphs
by: Zhu, Kerui, et al.
Published: (2024)
by: Zhu, Kerui, et al.
Published: (2024)
Instructive Decoding: Instruction-Tuned Large Language Models are Self-Refiner from Noisy Instructions
by: Kim, Taehyeon, et al.
Published: (2023)
by: Kim, Taehyeon, et al.
Published: (2023)
One-Topic-Doesn't-Fit-All: Transcreating Reading Comprehension Test for Personalized Learning
by: Han, Jieun, et al.
Published: (2025)
by: Han, Jieun, et al.
Published: (2025)
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
by: Barančíková, Petra, et al.
Published: (2025)
by: Barančíková, Petra, et al.
Published: (2025)
Improving Large Language Models with Concept-Aware Fine-Tuning
by: Chen, Michael K., et al.
Published: (2025)
by: Chen, Michael K., et al.
Published: (2025)
Safety-Aware Fine-Tuning of Large Language Models
by: Choi, Hyeong Kyu, et al.
Published: (2024)
by: Choi, Hyeong Kyu, et al.
Published: (2024)
Similar Items
-
PANDA: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation
by: Zhong, Qihuang, et al.
Published: (2022) -
E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation
by: Zhong, Qihuang, et al.
Published: (2022) -
Revisiting Knowledge Distillation for Autoregressive Language Models
by: Zhong, Qihuang, et al.
Published: (2024) -
KaFT: Knowledge-aware Fine-tuning for Boosting LLMs' Domain-specific Question-Answering Performance
by: Zhong, Qihuang, et al.
Published: (2025) -
Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
by: Zhong, Qihuang, et al.
Published: (2026)