CONSCENDI: A Contrastive and Scenario-Guided Distillation Approach to Guardrail Models for Virtual Assistants
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Albert Yu, Nair, Varun, Schumacher, Elliot, Kannan, Anitha |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rare Disease Differential Diagnosis with Large Language Models at Scale: From Abdominal Actinomycosis to Wilson's Disease
by: Schumacher, Elliot, et al.
Published: (2025)
by: Schumacher, Elliot, et al.
Published: (2025)
Extrinsically-Focused Evaluation of Omissions in Medical Summarization
by: Schumacher, Elliot, et al.
Published: (2023)
by: Schumacher, Elliot, et al.
Published: (2023)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
by: Phan, Phuc, et al.
Published: (2024)
by: Phan, Phuc, et al.
Published: (2024)
Evolutionary Contrastive Distillation for Language Model Alignment
by: Katz-Samuels, Julian, et al.
Published: (2024)
by: Katz-Samuels, Julian, et al.
Published: (2024)
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
by: Ko, Jongwoo, et al.
Published: (2025)
by: Ko, Jongwoo, et al.
Published: (2025)
Language-Guided World Models: A Model-Based Approach to AI Control
by: Zhang, Alex, et al.
Published: (2024)
by: Zhang, Alex, et al.
Published: (2024)
Jill Watson: A Virtual Teaching Assistant powered by ChatGPT
by: Taneja, Karan, et al.
Published: (2024)
by: Taneja, Karan, et al.
Published: (2024)
Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
by: Neill, James O', et al.
Published: (2025)
by: Neill, James O', et al.
Published: (2025)
Noise Injection Systemically Degrades Large Language Model Safety Guardrails
by: Shahani, Prithviraj Singh, et al.
Published: (2025)
by: Shahani, Prithviraj Singh, et al.
Published: (2025)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
SafePred: A Predictive Guardrail for Computer-Using Agents via World Models
by: Chen, Yurun, et al.
Published: (2026)
by: Chen, Yurun, et al.
Published: (2026)
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)
by: Elesedy, Hayder, et al.
Published: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
by: Oh, Sejoon, et al.
Published: (2024)
by: Oh, Sejoon, et al.
Published: (2024)
Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
by: Gueorguieva, Anna-Maria, et al.
Published: (2025)
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
by: Yuan, Zhuowen, et al.
Published: (2024)
by: Yuan, Zhuowen, et al.
Published: (2024)
BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate
by: Mazza, Arnon, et al.
Published: (2026)
by: Mazza, Arnon, et al.
Published: (2026)
Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
by: Ilin, Aleksei, et al.
Published: (2025)
by: Ilin, Aleksei, et al.
Published: (2025)
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
by: Krishna, Kundan, et al.
Published: (2025)
by: Krishna, Kundan, et al.
Published: (2025)
LLM Pruning and Distillation in Practice: The Minitron Approach
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2024)
by: Sreenivas, Sharath Turuvekere, et al.
Published: (2024)
CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
by: Li, Xiaoya, et al.
Published: (2025)
by: Li, Xiaoya, et al.
Published: (2025)
Evaluating Large Language Models Using Contrast Sets: An Experimental Approach
by: Sanwal, Manish
Published: (2024)
by: Sanwal, Manish
Published: (2024)
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
by: Che, Xinyu, et al.
Published: (2026)
by: Che, Xinyu, et al.
Published: (2026)
Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
by: Padarha, Shreyansh
Published: (2025)
by: Padarha, Shreyansh
Published: (2025)
Large Continual Instruction Assistant
by: Qiao, Jingyang, et al.
Published: (2024)
by: Qiao, Jingyang, et al.
Published: (2024)
Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
by: Yu, Fengming, et al.
Published: (2025)
by: Yu, Fengming, et al.
Published: (2025)
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
by: Järviniemi, Olli, et al.
Published: (2024)
by: Järviniemi, Olli, et al.
Published: (2024)
ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models
by: Thonet, Thibaut, et al.
Published: (2024)
by: Thonet, Thibaut, et al.
Published: (2024)
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
by: Rao, Varun, et al.
Published: (2025)
by: Rao, Varun, et al.
Published: (2025)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
by: Wen, Xiaofei, et al.
Published: (2025)
by: Wen, Xiaofei, et al.
Published: (2025)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
by: Zhang, Songming, et al.
Published: (2025)
by: Zhang, Songming, et al.
Published: (2025)
DVM: A Bytecode Virtual Machine Approach for Dynamic Tensor Computation
by: Fang, Jingzhi, et al.
Published: (2026)
by: Fang, Jingzhi, et al.
Published: (2026)
On Teacher Hacking in Language Model Distillation
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Self-Distillation for Multi-Token Prediction
by: Zhao, Guoliang, et al.
Published: (2026)
by: Zhao, Guoliang, et al.
Published: (2026)
A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
by: Khan, Sheraz, et al.
Published: (2025)
by: Khan, Sheraz, et al.
Published: (2025)
Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search
by: Liu, Max, et al.
Published: (2024)
by: Liu, Max, et al.
Published: (2024)
LLaSA: Large Language and Structured Data Assistant
by: Xu, Yao, et al.
Published: (2024)
by: Xu, Yao, et al.
Published: (2024)
Needle in the Haystack for Memory Based Large Language Models
by: Nelson, Elliot, et al.
Published: (2024)
by: Nelson, Elliot, et al.
Published: (2024)
Similar Items
-
Rare Disease Differential Diagnosis with Large Language Models at Scale: From Abdominal Actinomycosis to Wilson's Disease
by: Schumacher, Elliot, et al.
Published: (2025) -
Extrinsically-Focused Evaluation of Omissions in Medical Summarization
by: Schumacher, Elliot, et al.
Published: (2023) -
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
by: Phan, Phuc, et al.
Published: (2024) -
Evolutionary Contrastive Distillation for Language Model Alignment
by: Katz-Samuels, Julian, et al.
Published: (2024) -
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
by: Ko, Jongwoo, et al.
Published: (2025)