SLM as Guardian: Pioneering AI Safety with Small Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kwon, Ohjoon, Jeon, Donghyeon, Choi, Nayoung, Cho, Gyu-Hwung, Kim, Changbong, Lee, Hyunwoo, Kang, Inho, Kim, Sun, Park, Taiwoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Taxonomy and Analysis of Sensitive User Queries in Generative AI Search
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024)
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024)
QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines
von: Kwon, Ohjoon, et al.
Veröffentlicht: (2025)
von: Kwon, Ohjoon, et al.
Veröffentlicht: (2025)
ZeroDL: Zero-shot Distribution Learning for Text Clustering via Large Language Models
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024)
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024)
From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines
von: Lee, Sunkyung, et al.
Veröffentlicht: (2026)
von: Lee, Sunkyung, et al.
Veröffentlicht: (2026)
Hierarchical Multi-Persona Induction from User Behavioral Logs: Learning Evidence-Grounded and Truthful Personas
von: Choi, Nayoung, et al.
Veröffentlicht: (2026)
von: Choi, Nayoung, et al.
Veröffentlicht: (2026)
SLM-Based Agentic AI with P-C-G: Optimized for Korean Tool Use
von: Jeon, Changhyun, et al.
Veröffentlicht: (2025)
von: Jeon, Changhyun, et al.
Veröffentlicht: (2025)
ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models
von: Yoon, Junho, et al.
Veröffentlicht: (2025)
von: Yoon, Junho, et al.
Veröffentlicht: (2025)
AcuRank: Uncertainty-Aware Adaptive Computation for Listwise Reranking
von: Yoon, Soyoung, et al.
Veröffentlicht: (2025)
von: Yoon, Soyoung, et al.
Veröffentlicht: (2025)
RRADistill: Distilling LLMs' Passage Ranking Ability for Long-Tail Queries Document Re-Ranking on a Search Engine
von: Choi, Nayoung, et al.
Veröffentlicht: (2024)
von: Choi, Nayoung, et al.
Veröffentlicht: (2024)
Diffusion Prior-Based Amortized Variational Inference for Noisy Inverse Problems
von: Lee, Sojin, et al.
Veröffentlicht: (2024)
von: Lee, Sojin, et al.
Veröffentlicht: (2024)
Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval
von: Song, Jonghyun, et al.
Veröffentlicht: (2025)
von: Song, Jonghyun, et al.
Veröffentlicht: (2025)
Stochastic Conditional Diffusion Models for Robust Semantic Image Synthesis
von: Ko, Juyeon, et al.
Veröffentlicht: (2024)
von: Ko, Juyeon, et al.
Veröffentlicht: (2024)
The thermodynamic uncertainty relation of a quantum-mechanically coupled two-qubit system
von: Cho, Kwang Hyun, et al.
Veröffentlicht: (2025)
von: Cho, Kwang Hyun, et al.
Veröffentlicht: (2025)
Machine Learning‐Based Regression Modeling Approach to Predict the Carbon Content in Molten Steel of Electric Arc Furnace
von: Hyukjun Ha, et al.
Veröffentlicht: (2025)
von: Hyukjun Ha, et al.
Veröffentlicht: (2025)
Error as Signal: Stiffness-Aware Diffusion Sampling via Embedded Runge-Kutta Guidance
von: Kong, Inho, et al.
Veröffentlicht: (2026)
von: Kong, Inho, et al.
Veröffentlicht: (2026)
A Survey on Integration of Large Language Models with Intelligent Robots
von: Kim, Yeseung, et al.
Veröffentlicht: (2024)
von: Kim, Yeseung, et al.
Veröffentlicht: (2024)
AI‐based dementia risk prediction using voice digial biomarkers: A web‐based platform
von: Hyunwoong Ko, et al.
Veröffentlicht: (2024)
von: Hyunwoong Ko, et al.
Veröffentlicht: (2024)
Cell‐in‐Shell Metacells in Single‐Cell Nanoencapsulation
von: Duc Tai Nguyen, et al.
Veröffentlicht: (2026)
von: Duc Tai Nguyen, et al.
Veröffentlicht: (2026)
Zero-Shot Multi-Hop Question Answering via Monte-Carlo Tree Search with Large Language Models
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
von: Lee, Seongmin, et al.
Veröffentlicht: (2024)
Development of an Agentic AI Model for NGS Downstream Analysis Targeting Researchers with Limited Biological Background
von: Lee, Donghyeon, et al.
Veröffentlicht: (2025)
von: Lee, Donghyeon, et al.
Veröffentlicht: (2025)
Pioneering Carboxylated Zirconium Oxo Cluster Resist for Precision Nanoscale Patterning
von: Seong‐Ji Ha, et al.
Veröffentlicht: (2025)
von: Seong‐Ji Ha, et al.
Veröffentlicht: (2025)
Restricted weak type endpoint estimate for the spherical maximal operators on the Heisenberg group
von: Jeon, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Jeon, Hyunwoo, et al.
Veröffentlicht: (2025)
Beyond Economic Independence: A Longitudinal Study of ‘Self‐Resilience’ Trajectories Among Korean Care Leavers
von: Nayoung Kim
Veröffentlicht: (2026)
von: Nayoung Kim
Veröffentlicht: (2026)
Wearable Frontal EEG for Screening Cognitive Impairment
von: Nayoung Ryoo, et al.
Veröffentlicht: (2024)
von: Nayoung Ryoo, et al.
Veröffentlicht: (2024)
Correlation-driven tunability of altermagnetism in RuO$_2$
von: Park, Ina, et al.
Veröffentlicht: (2026)
von: Park, Ina, et al.
Veröffentlicht: (2026)
Synergistic Effect of Multilayered Alginate/Poly(SBMA) Coatings on Marine Antifouling Property
von: Jinwoo Lee, et al.
Veröffentlicht: (2024)
von: Jinwoo Lee, et al.
Veröffentlicht: (2024)
Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2025)
LINGO-Space: Language-Conditioned Incremental Grounding for Space
von: Kim, Dohyun, et al.
Veröffentlicht: (2024)
von: Kim, Dohyun, et al.
Veröffentlicht: (2024)
On the normalized local volume of a non-closed point
von: Kim, Donghyeon
Veröffentlicht: (2026)
von: Kim, Donghyeon
Veröffentlicht: (2026)
Asymptotically flat divisors and strongly $F$-regular type varieties
von: Kim, Donghyeon
Veröffentlicht: (2025)
von: Kim, Donghyeon
Veröffentlicht: (2025)
On a cohomological property of the center of a resolution
von: Kim, Donghyeon
Veröffentlicht: (2022)
von: Kim, Donghyeon
Veröffentlicht: (2022)
The number of smooth varieties in an MMP on a 3-fold of Fano type
von: Kim, Donghyeon
Veröffentlicht: (2025)
von: Kim, Donghyeon
Veröffentlicht: (2025)
On diminished multiplier ideal and the termination of flips
von: Kim, Donghyeon
Veröffentlicht: (2024)
von: Kim, Donghyeon
Veröffentlicht: (2024)
On volumes and the generic invariance of Fano type varieties
von: Kim, Donghyeon
Veröffentlicht: (2025)
von: Kim, Donghyeon
Veröffentlicht: (2025)
Rational singularities and $q$-birational morphism
von: Kim, Donghyeon
Veröffentlicht: (2023)
von: Kim, Donghyeon
Veröffentlicht: (2023)
Attention-aware Semantic Communications for Collaborative Inference
von: Im, Jiwoong, et al.
Veröffentlicht: (2024)
von: Im, Jiwoong, et al.
Veröffentlicht: (2024)
Citrus sunki Peel Extract, Containing Nobiletin and Tangeretin, Enhances Proliferation and Differentiation in 3T3-L1 Adipocytes
von: Sehyuk Oh, et al.
Veröffentlicht: (2024)
von: Sehyuk Oh, et al.
Veröffentlicht: (2024)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
The impact of corporate social responsibility on employee burnout: The crucial role of work overload
von: Byung‐Jik Kim, et al.
Veröffentlicht: (2024)
von: Byung‐Jik Kim, et al.
Veröffentlicht: (2024)
Robust Weight Initialization for Tanh Neural Networks with Fixed Point Analysis
von: Lee, Hyunwoo, et al.
Veröffentlicht: (2024)
von: Lee, Hyunwoo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Taxonomy and Analysis of Sensitive User Queries in Generative AI Search
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024) -
QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines
von: Kwon, Ohjoon, et al.
Veröffentlicht: (2025) -
ZeroDL: Zero-shot Distribution Learning for Text Clustering via Large Language Models
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024) -
From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines
von: Lee, Sunkyung, et al.
Veröffentlicht: (2026) -
Hierarchical Multi-Persona Induction from User Behavioral Logs: Learning Evidence-Grounded and Truthful Personas
von: Choi, Nayoung, et al.
Veröffentlicht: (2026)