SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Seanie, Lee, Dong Bok, Wagner, Dominik, Kang, Minki, Seong, Haebin, Bocklet, Tobias, Lee, Juho, Hwang, Sung Ju |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
by: Lee, Seanie, et al.
Published: (2025)
by: Lee, Seanie, et al.
Published: (2025)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Distilling LLM Agent into Small Models with Retrieval and Code Tools
by: Kang, Minki, et al.
Published: (2025)
by: Kang, Minki, et al.
Published: (2025)
Rethinking Reward Models for Multi-Domain Test-Time Scaling
by: Lee, Dong Bok, et al.
Published: (2025)
by: Lee, Dong Bok, et al.
Published: (2025)
Self-Supervised Dataset Distillation for Transfer Learning
by: Lee, Dong Bok, et al.
Published: (2023)
by: Lee, Dong Bok, et al.
Published: (2023)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2025)
by: Lee, Youngwan, et al.
Published: (2025)
Set-based Meta-Interpolation for Few-Task Meta-Learning
by: Lee, Seanie, et al.
Published: (2022)
by: Lee, Seanie, et al.
Published: (2022)
THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
by: Lee, Seanie, et al.
Published: (2026)
by: Lee, Seanie, et al.
Published: (2026)
Optimized Speculative Sampling for GPU Hardware Accelerators
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
DiffusionNAG: Predictor-guided Neural Architecture Generation with Diffusion Models
by: An, Sohyun, et al.
Published: (2023)
by: An, Sohyun, et al.
Published: (2023)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
by: Lee, Isack, et al.
Published: (2024)
by: Lee, Isack, et al.
Published: (2024)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Language Models
by: Kang, Minki, et al.
Published: (2024)
by: Kang, Minki, et al.
Published: (2024)
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Drug Discovery with Dynamic Goal-aware Fragments
by: Lee, Seul, et al.
Published: (2023)
by: Lee, Seul, et al.
Published: (2023)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
by: Wagner, Dominik, et al.
Published: (2025)
by: Wagner, Dominik, et al.
Published: (2025)
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
FedRand: Enhancing Privacy in Federated Learning with Randomized LoRA Subparameter Updates
by: Park, Sangwoo, et al.
Published: (2025)
by: Park, Sangwoo, et al.
Published: (2025)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
by: Lee, Dong Bok, et al.
Published: (2025)
by: Lee, Dong Bok, et al.
Published: (2025)
Reliable Decision Making via Calibration Oriented Retrieval Augmented Generation
by: Jang, Chaeyun, et al.
Published: (2024)
by: Jang, Chaeyun, et al.
Published: (2024)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
by: Wagner, Dominik, et al.
Published: (2023)
by: Wagner, Dominik, et al.
Published: (2023)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
Cost-Sensitive Multi-Fidelity Bayesian Optimization with Transfer of Learning Curve Extrapolation
by: Lee, Dong Bok, et al.
Published: (2024)
by: Lee, Dong Bok, et al.
Published: (2024)
Concept-skill Transferability-based Data Selection for Large Vision-Language Models
by: Lee, Jaewoo, et al.
Published: (2024)
by: Lee, Jaewoo, et al.
Published: (2024)
Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks
by: Lee, Dongwoo, et al.
Published: (2025)
by: Lee, Dongwoo, et al.
Published: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
by: Willette, Jeffrey, et al.
Published: (2025)
by: Willette, Jeffrey, et al.
Published: (2025)
KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis
by: Lee, Youngwan, et al.
Published: (2023)
by: Lee, Youngwan, et al.
Published: (2023)
GFlowPO: Generative Flow Network as a Language Model Prompt Optimizer
by: Cho, Junmo, et al.
Published: (2026)
by: Cho, Junmo, et al.
Published: (2026)
Efficient Long Context Language Model Retrieval with Compression
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
by: Park, Sangwoo, et al.
Published: (2026)
by: Park, Sangwoo, et al.
Published: (2026)
Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
by: Engert, Natalie, et al.
Published: (2026)
by: Engert, Natalie, et al.
Published: (2026)
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
by: Baumann, Ilja, et al.
Published: (2025)
by: Baumann, Ilja, et al.
Published: (2025)
Theory of Slidetronics in Ferroelectric van der Waals Layers
by: Lee, Byeoksong, et al.
Published: (2025)
by: Lee, Byeoksong, et al.
Published: (2025)
Seismic Safety Evaluation of a Full‐Size Reinforced Concrete Frame Infilled With Precast Modular Reinforced Blocks Using Pseudo‐Dynamic Testing and Nonlinear Dynamic Finite Element Analysis
by: Ju‐Seong Jung, et al.
Published: (2024)
by: Ju‐Seong Jung, et al.
Published: (2024)
Learning diverse attacks on large language models for robust red-teaming and safety tuning
by: Lee, Seanie, et al.
Published: (2024)
by: Lee, Seanie, et al.
Published: (2024)
Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model
by: Kim, Dongki, et al.
Published: (2025)
by: Kim, Dongki, et al.
Published: (2025)
SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
Similar Items
-
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
by: Lee, Seanie, et al.
Published: (2025) -
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
by: Lee, Seanie, et al.
Published: (2024) -
Distilling LLM Agent into Small Models with Retrieval and Code Tools
by: Kang, Minki, et al.
Published: (2025) -
Rethinking Reward Models for Multi-Domain Test-Time Scaling
by: Lee, Dong Bok, et al.
Published: (2025) -
Self-Supervised Dataset Distillation for Transfer Learning
by: Lee, Dong Bok, et al.
Published: (2023)