Exploring and Developing a Pre-Model Safeguard with Draft Models
Fuente:
Zenodo
Saved in:
| Main Authors: | Cai, Hongyu, Arunasalam, Arjun, Liang, Yiming, Bianchi, Antonio, Celik, Z. Berkay |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2026
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring and Developing a Pre-Model Safeguard with Draft Models
by: Cai, Hongyu, et al.
Published: (2026)
by: Cai, Hongyu, et al.
Published: (2026)
Rethinking How to Evaluate Language Model Jailbreak
by: Cai, Hongyu, et al.
Published: (2024)
by: Cai, Hongyu, et al.
Published: (2024)
International Students and Scams: At Risk Abroad
by: Zhang, Katherine, et al.
Published: (2025)
by: Zhang, Katherine, et al.
Published: (2025)
Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday Tasks
by: Arunasalam, Arjun, et al.
Published: (2025)
by: Arunasalam, Arjun, et al.
Published: (2025)
A Progressive Transformer for Unifying Binary Code Embedding and Knowledge Transfer
by: Lu, Hanxiao, et al.
Published: (2024)
by: Lu, Hanxiao, et al.
Published: (2024)
Investigating the Impact of Dark Patterns on LLM-Based Web Agents
by: Ersoy, Devin, et al.
Published: (2025)
by: Ersoy, Devin, et al.
Published: (2025)
RoboJailBench: Benchmarking Adversarial Attacks and Defenses in Embodied Robotic Agents
by: Yeke, Doguhuan, et al.
Published: (2026)
by: Yeke, Doguhuan, et al.
Published: (2026)
LM-Scout: Analyzing the Security of Language Model Integration in Android Apps
by: Ibrahim, Muhammad, et al.
Published: (2025)
by: Ibrahim, Muhammad, et al.
Published: (2025)
Understanding Users' Security and Privacy Concerns and Attitudes Towards Conversational AI Platforms
by: Ali, Mutahar, et al.
Published: (2025)
by: Ali, Mutahar, et al.
Published: (2025)
Enhancing LLM-based Autonomous Driving Agents to Mitigate Perception Attacks
by: Song, Ruoyu, et al.
Published: (2024)
by: Song, Ruoyu, et al.
Published: (2024)
Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems
by: Allegrini, Edoardo, et al.
Published: (2025)
by: Allegrini, Edoardo, et al.
Published: (2025)
D-LiFT: Improving LLM-based Decompiler Backend via Code Quality-driven Fine-tuning
by: Zou, Muqi, et al.
Published: (2025)
by: Zou, Muqi, et al.
Published: (2025)
STARS: Synchronous Token Alignment for Robust Supervision in Large Language Models
by: Quamar, Mohammad Atif, et al.
Published: (2025)
by: Quamar, Mohammad Atif, et al.
Published: (2025)
MAUI: Reconstructing Private Client Data in Federated Transfer Learning
by: Dabholkar, Ahaan, et al.
Published: (2025)
by: Dabholkar, Ahaan, et al.
Published: (2025)
Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models
by: Liang, Haoyu, et al.
Published: (2025)
by: Liang, Haoyu, et al.
Published: (2025)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
by: Lin, Chun-Hsien, et al.
Published: (2024)
by: Lin, Chun-Hsien, et al.
Published: (2024)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
by: Lei, Haodi, et al.
Published: (2026)
by: Lei, Haodi, et al.
Published: (2026)
Side-channel Inference of User Activities in AR/VR Using GPU Profiling
by: Son, Seonghun, et al.
Published: (2025)
by: Son, Seonghun, et al.
Published: (2025)
Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
by: Farn, Hua, et al.
Published: (2024)
by: Farn, Hua, et al.
Published: (2024)
Exploring and Expanding Secondary Findings Through Exome Sequencing in the Turkish Population
by: Mehmet Berkay Akcan, et al.
Published: (2025)
by: Mehmet Berkay Akcan, et al.
Published: (2025)
Adaptive Blockwise Search: Inference-Time Alignment for Large Language Models
by: Quamar, Mohammad Atif, et al.
Published: (2025)
by: Quamar, Mohammad Atif, et al.
Published: (2025)
The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents
by: Yeke, Doguhan, et al.
Published: (2026)
by: Yeke, Doguhan, et al.
Published: (2026)
Exploring and Improving Drafts in Blockwise Parallel Decoding
by: Kim, Taehyeon, et al.
Published: (2024)
by: Kim, Taehyeon, et al.
Published: (2024)
A Survey on Pre-Trained Diffusion Model Distillations
by: Fan, Xuhui, et al.
Published: (2025)
by: Fan, Xuhui, et al.
Published: (2025)
AdaFortiTran: An Adaptive Transformer Model for Robust OFDM Channel Estimation
by: Guler, Berkay, et al.
Published: (2025)
by: Guler, Berkay, et al.
Published: (2025)
Discussion on the Relation Between Carbon Emissions and Solid Waste Management to Develop a Sustainable Business Model
by: Cemre Avşar, et al.
Published: (2025)
by: Cemre Avşar, et al.
Published: (2025)
Physical ID-Transfer Attacks against Multi-Object Tracking via Adversarial Trajectory
by: Wang, Chenyi, et al.
Published: (2025)
by: Wang, Chenyi, et al.
Published: (2025)
Exploring genetic variants in congenital monosaccharide‐disaccharide metabolism: Carrier ratios and phenotypic insights
by: Mehmet Berkay Akcan, et al.
Published: (2024)
by: Mehmet Berkay Akcan, et al.
Published: (2024)
Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
by: Lei, Haodong, et al.
Published: (2025)
by: Lei, Haodong, et al.
Published: (2025)
‘Safeguarding’, a key dispositif of the ICH convention
by: Antonio Arantes
Published: (2019)
by: Antonio Arantes
Published: (2019)
DEER: Draft with Diffusion, Verify with Autoregressive Models
by: Cheng, Zicong, et al.
Published: (2025)
by: Cheng, Zicong, et al.
Published: (2025)
NFL Draft Modelling: Loss Functional Analysis
by: Grandhisiri, Tanmay
Published: (2025)
by: Grandhisiri, Tanmay
Published: (2025)
Development of a MATLAB/Simulink Based Performance and Simulation Program for Fuel Cell Hybrid Electric Wheeled and Tracked Vehicle Powertrains
by: Berkay Açık, et al.
Published: (2026)
by: Berkay Açık, et al.
Published: (2026)
Exploring Generalizable Pre-training for Real-world Change Detection via Geometric Estimation
by: Zhao, Yitao, et al.
Published: (2025)
by: Zhao, Yitao, et al.
Published: (2025)
Model Capability Assessment and Safeguards for Biological Weaponization
by: Richter, Michael
Published: (2026)
by: Richter, Michael
Published: (2026)
On Prompt-Driven Safeguarding for Large Language Models
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
Similar Items
-
Exploring and Developing a Pre-Model Safeguard with Draft Models
by: Cai, Hongyu, et al.
Published: (2026) -
Rethinking How to Evaluate Language Model Jailbreak
by: Cai, Hongyu, et al.
Published: (2024) -
International Students and Scams: At Risk Abroad
by: Zhang, Katherine, et al.
Published: (2025) -
Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday Tasks
by: Arunasalam, Arjun, et al.
Published: (2025) -
A Progressive Transformer for Unifying Binary Code Embedding and Knowledge Transfer
by: Lu, Hanxiao, et al.
Published: (2024)