HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Seanie, Seong, Haebin, Lee, Dong Bok, Kang, Minki, Chen, Xiaoyin, Wagner, Dominik, Bengio, Yoshua, Lee, Juho, Hwang, Sung Ju |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
by: Lee, Seanie, et al.
Published: (2025)
by: Lee, Seanie, et al.
Published: (2025)
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
by: Lee, Seanie, et al.
Published: (2025)
by: Lee, Seanie, et al.
Published: (2025)
Self-Supervised Dataset Distillation for Transfer Learning
by: Lee, Dong Bok, et al.
Published: (2023)
by: Lee, Dong Bok, et al.
Published: (2023)
Distilling LLM Agent into Small Models with Retrieval and Code Tools
by: Kang, Minki, et al.
Published: (2025)
by: Kang, Minki, et al.
Published: (2025)
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
Set-based Meta-Interpolation for Few-Task Meta-Learning
by: Lee, Seanie, et al.
Published: (2022)
by: Lee, Seanie, et al.
Published: (2022)
Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language Models
by: Kang, Minki, et al.
Published: (2024)
by: Kang, Minki, et al.
Published: (2024)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Drug Discovery with Dynamic Goal-aware Fragments
by: Lee, Seul, et al.
Published: (2023)
by: Lee, Seul, et al.
Published: (2023)
THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
by: Lee, Seanie, et al.
Published: (2026)
by: Lee, Seanie, et al.
Published: (2026)
Reliable Decision Making via Calibration Oriented Retrieval Augmented Generation
by: Jang, Chaeyun, et al.
Published: (2024)
by: Jang, Chaeyun, et al.
Published: (2024)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
by: Lee, Isack, et al.
Published: (2024)
by: Lee, Isack, et al.
Published: (2024)
DiffusionNAG: Predictor-guided Neural Architecture Generation with Diffusion Models
by: An, Sohyun, et al.
Published: (2023)
by: An, Sohyun, et al.
Published: (2023)
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Rethinking Reward Models for Multi-Domain Test-Time Scaling
by: Lee, Dong Bok, et al.
Published: (2025)
by: Lee, Dong Bok, et al.
Published: (2025)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2025)
by: Lee, Youngwan, et al.
Published: (2025)
FedRand: Enhancing Privacy in Federated Learning with Randomized LoRA Subparameter Updates
by: Park, Sangwoo, et al.
Published: (2025)
by: Park, Sangwoo, et al.
Published: (2025)
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
by: Park, Sangwoo, et al.
Published: (2026)
by: Park, Sangwoo, et al.
Published: (2026)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
Cost-Sensitive Multi-Fidelity Bayesian Optimization with Transfer of Learning Curve Extrapolation
by: Lee, Dong Bok, et al.
Published: (2024)
by: Lee, Dong Bok, et al.
Published: (2024)
Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks
by: Lee, Dongwoo, et al.
Published: (2025)
by: Lee, Dongwoo, et al.
Published: (2025)
KTRL+F: Knowledge-Augmented In-Document Search
by: Oh, Hanseok, et al.
Published: (2023)
by: Oh, Hanseok, et al.
Published: (2023)
OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
by: Baek, Jinheon, et al.
Published: (2026)
by: Baek, Jinheon, et al.
Published: (2026)
Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
by: Lee, Dong Bok, et al.
Published: (2025)
by: Lee, Dong Bok, et al.
Published: (2025)
Theory of Slidetronics in Ferroelectric van der Waals Layers
by: Lee, Byeoksong, et al.
Published: (2025)
by: Lee, Byeoksong, et al.
Published: (2025)
Efficient Causal Graph Discovery Using Large Language Models
by: Jiralerspong, Thomas, et al.
Published: (2024)
by: Jiralerspong, Thomas, et al.
Published: (2024)
Seismic Safety Evaluation of a Full‐Size Reinforced Concrete Frame Infilled With Precast Modular Reinforced Blocks Using Pseudo‐Dynamic Testing and Nonlinear Dynamic Finite Element Analysis
by: Ju‐Seong Jung, et al.
Published: (2024)
by: Ju‐Seong Jung, et al.
Published: (2024)
Atomically‐Thin Holey 2D Nanosheets of Defect‐Engineered MoN–Mo5N6 Composites as Effective Hybridization Matrices (Small 9/2024)
by: Jihyeong Lee, et al.
Published: (2024)
by: Jihyeong Lee, et al.
Published: (2024)
Optimized Speculative Sampling for GPU Hardware Accelerators
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
Machine learning and information theory concepts towards an AI Mathematician
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
PREPING: Building Agent Memory without Tasks
by: Choi, Yumin, et al.
Published: (2026)
by: Choi, Yumin, et al.
Published: (2026)
Filling the Gaps: Selective Knowledge Augmentation for LLM Recommenders
by: Lee, Jaehyun, et al.
Published: (2026)
by: Lee, Jaehyun, et al.
Published: (2026)
GFlowPO: Generative Flow Network as a Language Model Prompt Optimizer
by: Cho, Junmo, et al.
Published: (2026)
by: Cho, Junmo, et al.
Published: (2026)
Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
by: Kang, Minki, et al.
Published: (2026)
by: Kang, Minki, et al.
Published: (2026)
Baking Symmetry into GFlowNets
by: Ma, George, et al.
Published: (2024)
by: Ma, George, et al.
Published: (2024)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
by: Wagner, Dominik, et al.
Published: (2025)
by: Wagner, Dominik, et al.
Published: (2025)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
by: Williams-King, David, et al.
Published: (2025)
by: Williams-King, David, et al.
Published: (2025)
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
by: Bartoldson, Brian, et al.
Published: (2025)
by: Bartoldson, Brian, et al.
Published: (2025)
Similar Items
-
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
by: Lee, Seanie, et al.
Published: (2025) -
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
by: Lee, Seanie, et al.
Published: (2025) -
Self-Supervised Dataset Distillation for Transfer Learning
by: Lee, Dong Bok, et al.
Published: (2023) -
Distilling LLM Agent into Small Models with Retrieval and Code Tools
by: Kang, Minki, et al.
Published: (2025) -
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
by: Kang, Seongjae, et al.
Published: (2025)