Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arif, Huzaifa, Murugesan, Keerthiram, Ko, Ching-Yun, Chen, Pin-Yu, Das, Payel, Gittens, Alex |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PEEL the Layers and Find Yourself: Revisiting Inference-time Data Leakage for Residual Neural Networks
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
Forecasting Fails: Unveiling Evasion Attacks in Weather Prediction Models
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025)
Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models
von: Chen, Pin-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Pin-Yu, et al.
Veröffentlicht: (2025)
vLLM Hook v0: A Plug-in for Programming Model Internals on vLLM
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2026)
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2026)
STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language Models
von: Basavatia, Shreyas, et al.
Veröffentlicht: (2024)
von: Basavatia, Shreyas, et al.
Veröffentlicht: (2024)
Large Language Models can be Strong Self-Detoxifiers
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2024)
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2024)
SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection
von: Shen, Han, et al.
Veröffentlicht: (2024)
von: Shen, Han, et al.
Veröffentlicht: (2024)
DS FedProxGrad: Asymptotic Stationarity Without Noise Floor in Fair Federated Learning
von: Arif, Huzaifa
Veröffentlicht: (2025)
von: Arif, Huzaifa
Veröffentlicht: (2025)
Can Memory-Augmented Language Models Generalize on Reasoning-in-a-Haystack Tasks?
von: Das, Payel, et al.
Veröffentlicht: (2025)
von: Das, Payel, et al.
Veröffentlicht: (2025)
On the Effects of Fine-tuning Language Models for Text-Based Reinforcement Learning
von: Gruppi, Mauricio, et al.
Veröffentlicht: (2024)
von: Gruppi, Mauricio, et al.
Veröffentlicht: (2024)
Towards Aligning Language Models with Textual Feedback
von: Lloret, Saüc Abadal, et al.
Veröffentlicht: (2024)
von: Lloret, Saüc Abadal, et al.
Veröffentlicht: (2024)
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
von: Villa, Danielle, et al.
Veröffentlicht: (2025)
von: Villa, Danielle, et al.
Veröffentlicht: (2025)
Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
von: Xiong, Chen, et al.
Veröffentlicht: (2026)
von: Xiong, Chen, et al.
Veröffentlicht: (2026)
Iterative thresholding for non-linear learning in the strong $\varepsilon$-contamination model
von: Rathnashyam, Arvind, et al.
Veröffentlicht: (2024)
von: Rathnashyam, Arvind, et al.
Veröffentlicht: (2024)
Extending Model-x Framework to Missing Data
von: Koyuncu, Deniz, et al.
Veröffentlicht: (2022)
von: Koyuncu, Deniz, et al.
Veröffentlicht: (2022)
Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning
von: Lee, Yu-Ang, et al.
Veröffentlicht: (2026)
von: Lee, Yu-Ang, et al.
Veröffentlicht: (2026)
The Unlearning Mirage: A Dynamic Framework for Evaluating LLM Unlearning
von: Shah, Raj Sanjay, et al.
Veröffentlicht: (2026)
von: Shah, Raj Sanjay, et al.
Veröffentlicht: (2026)
Language Guided Exploration for RL Agents in Text Environments
von: Golchha, Hitesh, et al.
Veröffentlicht: (2024)
von: Golchha, Hitesh, et al.
Veröffentlicht: (2024)
EfficientLLM: Efficiency in Large Language Models
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2025)
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2025)
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
von: Peng, ShengYun, et al.
Veröffentlicht: (2024)
OjaKV: Context-Aware Online Low-Rank KV Cache Compression
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Understanding and Improving Training-Free AI-Generated Image Detections with Vision Foundation Models
von: Tsai, Chung-Ting, et al.
Veröffentlicht: (2024)
von: Tsai, Chung-Ting, et al.
Veröffentlicht: (2024)
Targeted Advertising on Social Networks Using Online Variational Tensor Regression
von: Idé, Tsuyoshi, et al.
Veröffentlicht: (2022)
von: Idé, Tsuyoshi, et al.
Veröffentlicht: (2022)
Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs
von: Thakkar, Megh, et al.
Veröffentlicht: (2024)
von: Thakkar, Megh, et al.
Veröffentlicht: (2024)
Defining and Evaluating Physical Safety for Large Language Models
von: Tang, Yung-Chen, et al.
Veröffentlicht: (2024)
von: Tang, Yung-Chen, et al.
Veröffentlicht: (2024)
Language Models Coupled with Metacognition Can Outperform Reasoning Models
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2025)
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2025)
Interpretable Graph-Language Modeling for Detecting Youth Illicit Drug Use
von: Li, Yiyang, et al.
Veröffentlicht: (2025)
von: Li, Yiyang, et al.
Veröffentlicht: (2025)
Needle in the Haystack for Memory Based Large Language Models
von: Nelson, Elliot, et al.
Veröffentlicht: (2024)
von: Nelson, Elliot, et al.
Veröffentlicht: (2024)
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value Retrieval
von: Yang, David H., et al.
Veröffentlicht: (2026)
von: Yang, David H., et al.
Veröffentlicht: (2026)
On the Prospects of Incorporating Large Language Models (LLMs) in Automated Planning and Scheduling (APS)
von: Pallagani, Vishal, et al.
Veröffentlicht: (2024)
von: Pallagani, Vishal, et al.
Veröffentlicht: (2024)
LongDA: Benchmarking LLM Agents for Long-Document Data Analysis
von: Li, Yiyang, et al.
Veröffentlicht: (2026)
von: Li, Yiyang, et al.
Veröffentlicht: (2026)
LLM-Enhanced Software Patch Localization
von: Yu, Jinhong, et al.
Veröffentlicht: (2024)
von: Yu, Jinhong, et al.
Veröffentlicht: (2024)
DistiLLM: Towards Streamlined Distillation for Large Language Models
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design
von: Neehal, Nafis, et al.
Veröffentlicht: (2024)
von: Neehal, Nafis, et al.
Veröffentlicht: (2024)
SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning
von: Zhang, Shuai, et al.
Veröffentlicht: (2024)
von: Zhang, Shuai, et al.
Veröffentlicht: (2024)
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing
von: Mukherjee, Arpan, et al.
Veröffentlicht: (2024)
von: Mukherjee, Arpan, et al.
Veröffentlicht: (2024)
Context Attribution with Multi-Armed Bandit Optimization
von: Pan, Deng, et al.
Veröffentlicht: (2025)
von: Pan, Deng, et al.
Veröffentlicht: (2025)
Larimar: Large Language Models with Episodic Memory Control
von: Das, Payel, et al.
Veröffentlicht: (2024)
von: Das, Payel, et al.
Veröffentlicht: (2024)
Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
von: Chen, Kejia, et al.
Veröffentlicht: (2025)
von: Chen, Kejia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PEEL the Layers and Find Yourself: Revisiting Inference-time Data Leakage for Residual Neural Networks
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025) -
Forecasting Fails: Unveiling Evasion Attacks in Weather Prediction Models
von: Arif, Huzaifa, et al.
Veröffentlicht: (2025) -
Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models
von: Chen, Pin-Yu, et al.
Veröffentlicht: (2025) -
vLLM Hook v0: A Plug-in for Programming Model Internals on vLLM
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2026) -
STARLING: Self-supervised Training of Text-based Reinforcement Learning Agent with Large Language Models
von: Basavatia, Shreyas, et al.
Veröffentlicht: (2024)