Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Addepalli, Sravanti, Varun, Yerram, Suggala, Arun, Shanmugam, Karthikeyan, Jain, Prateek |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Time-Reversal Provides Unsupervised Feedback to LLMs
von: Varun, Yerram, et al.
Veröffentlicht: (2024)
von: Varun, Yerram, et al.
Veröffentlicht: (2024)
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
von: Nair, Pranav Ajit, et al.
Veröffentlicht: (2024)
von: Nair, Pranav Ajit, et al.
Veröffentlicht: (2024)
Does Refusal Training in LLMs Generalize to the Past Tense?
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Tandem Transformers for Inference Efficient LLMs
von: S, Aishwarya P, et al.
Veröffentlicht: (2024)
von: S, Aishwarya P, et al.
Veröffentlicht: (2024)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
von: Thomas, Rohan Subramanian, et al.
Veröffentlicht: (2026)
von: Thomas, Rohan Subramanian, et al.
Veröffentlicht: (2026)
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
von: Li, Mingjie, et al.
Veröffentlicht: (2026)
von: Li, Mingjie, et al.
Veröffentlicht: (2026)
Semantic Mastery: Enhancing LLMs with Advanced Natural Language Understanding
von: Hariharan, Mohanakrishnan
Veröffentlicht: (2025)
von: Hariharan, Mohanakrishnan
Veröffentlicht: (2025)
Interleaved Gibbs Diffusion: Generating Discrete-Continuous Data with Implicit Constraints
von: Anil, Gautham Govind, et al.
Veröffentlicht: (2025)
von: Anil, Gautham Govind, et al.
Veröffentlicht: (2025)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
von: He, Qianxi, et al.
Veröffentlicht: (2025)
von: He, Qianxi, et al.
Veröffentlicht: (2025)
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
von: Ren, Qibing, et al.
Veröffentlicht: (2024)
von: Ren, Qibing, et al.
Veröffentlicht: (2024)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
von: Goru, Ritesh, et al.
Veröffentlicht: (2025)
von: Goru, Ritesh, et al.
Veröffentlicht: (2025)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
von: Ning, Xuefei, et al.
Veröffentlicht: (2023)
von: Ning, Xuefei, et al.
Veröffentlicht: (2023)
Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
von: Cai, Hanyu, et al.
Veröffentlicht: (2025)
von: Cai, Hanyu, et al.
Veröffentlicht: (2025)
Reasoning about concepts with LLMs: Inconsistencies abound
von: Uceda-Sosa, Rosario, et al.
Veröffentlicht: (2024)
von: Uceda-Sosa, Rosario, et al.
Veröffentlicht: (2024)
CPTuning: Contrastive Prompt Tuning for Generative Relation Extraction
von: Duan, Jiaxin, et al.
Veröffentlicht: (2025)
von: Duan, Jiaxin, et al.
Veröffentlicht: (2025)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
von: Tan, Daniel, et al.
Veröffentlicht: (2025)
von: Tan, Daniel, et al.
Veröffentlicht: (2025)
Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
von: Chhikara, Prateek
Veröffentlicht: (2025)
von: Chhikara, Prateek
Veröffentlicht: (2025)
Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
von: Si, Jianfeng, et al.
Veröffentlicht: (2025)
von: Si, Jianfeng, et al.
Veröffentlicht: (2025)
Generative AI for FFRDCs
von: Maiya, Arun S.
Veröffentlicht: (2025)
von: Maiya, Arun S.
Veröffentlicht: (2025)
Supervisory Prompt Training
von: Billa, Jean Ghislain, et al.
Veröffentlicht: (2024)
von: Billa, Jean Ghislain, et al.
Veröffentlicht: (2024)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
von: Yuan, Youliang, et al.
Veröffentlicht: (2024)
Prompt-Based One-Shot Exact Length-Controlled Generation with LLMs
von: Xie, Juncheng, et al.
Veröffentlicht: (2025)
von: Xie, Juncheng, et al.
Veröffentlicht: (2025)
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
von: Yuan, Yuan, et al.
Veröffentlicht: (2025)
von: Yuan, Yuan, et al.
Veröffentlicht: (2025)
Can LLMs Generate Visualizations with Dataless Prompts?
von: Coelho, Darius, et al.
Veröffentlicht: (2024)
von: Coelho, Darius, et al.
Veröffentlicht: (2024)
HanjaBridge: Resolving Semantic Ambiguity in Korean LLMs via Hanja-Augmented Pre-Training
von: Choi, Seungho
Veröffentlicht: (2025)
von: Choi, Seungho
Veröffentlicht: (2025)
A Lightweight Explainable Guardrail for Prompt Safety
von: Islam, Md Asiful, et al.
Veröffentlicht: (2026)
von: Islam, Md Asiful, et al.
Veröffentlicht: (2026)
LLMs can Compress LLMs: Adaptive Pruning by Agents
von: Kodathala, Sai Varun, et al.
Veröffentlicht: (2026)
von: Kodathala, Sai Varun, et al.
Veröffentlicht: (2026)
Zero-Shot Continuous Prompt Transfer: Generalizing Task Semantics Across Language Models
von: Wu, Zijun, et al.
Veröffentlicht: (2023)
von: Wu, Zijun, et al.
Veröffentlicht: (2023)
A Survey on Prompting Techniques in LLMs
von: Bhandari, Prabin
Veröffentlicht: (2023)
von: Bhandari, Prabin
Veröffentlicht: (2023)
Local Prompt Optimization
von: Jain, Yash, et al.
Veröffentlicht: (2025)
von: Jain, Yash, et al.
Veröffentlicht: (2025)
Diversity of Thought Improves Reasoning Abilities of LLMs
von: Naik, Ranjita, et al.
Veröffentlicht: (2023)
von: Naik, Ranjita, et al.
Veröffentlicht: (2023)
"When Data is Scarce, Prompt Smarter"... Approaches to Grammatical Error Correction in Low-Resource Settings
von: De, Somsubhra, et al.
Veröffentlicht: (2025)
von: De, Somsubhra, et al.
Veröffentlicht: (2025)
Prompting Implicit Discourse Relation Annotation
von: Yung, Frances, et al.
Veröffentlicht: (2024)
von: Yung, Frances, et al.
Veröffentlicht: (2024)
Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses
von: Peng, Kerui, et al.
Veröffentlicht: (2026)
von: Peng, Kerui, et al.
Veröffentlicht: (2026)
EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
von: Guo, Qingyan, et al.
Veröffentlicht: (2023)
von: Guo, Qingyan, et al.
Veröffentlicht: (2023)
PromptTailor: Multi-turn Intent-Aligned Prompt Synthesis for Lightweight LLMs
von: Xu, Yizhou, et al.
Veröffentlicht: (2025)
von: Xu, Yizhou, et al.
Veröffentlicht: (2025)
Graph-Augmented Relation Extraction Model with LLMs-Generated Support Document
von: Dong, Vicky, et al.
Veröffentlicht: (2024)
von: Dong, Vicky, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Time-Reversal Provides Unsupervised Feedback to LLMs
von: Varun, Yerram, et al.
Veröffentlicht: (2024) -
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
von: Nair, Pranav Ajit, et al.
Veröffentlicht: (2024) -
Does Refusal Training in LLMs Generalize to the Past Tense?
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024) -
Tandem Transformers for Inference Efficient LLMs
von: S, Aishwarya P, et al.
Veröffentlicht: (2024) -
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
von: Thomas, Rohan Subramanian, et al.
Veröffentlicht: (2026)