Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Minkyoung, Kim, Yunha, Seo, Hyeram, Choi, Heejung, Han, Jiye, Kee, Gaeun, Ko, Soyoung, Jung, HyoJe, Kim, Byeolhee, Kim, Young-Hak, Park, Sanghyun, Jun, Tae Joon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Clinical Efficiency through LLM: Discharge Note Generation for Cardiac Patients
by: Jung, HyoJe, et al.
Published: (2024)
by: Jung, HyoJe, et al.
Published: (2024)
Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering
by: Kim, Byeolhee, et al.
Published: (2026)
by: Kim, Byeolhee, et al.
Published: (2026)
InMD-X: Large Language Models for Internal Medicine Doctors
by: Gwon, Hansle, et al.
Published: (2024)
by: Gwon, Hansle, et al.
Published: (2024)
Multi-Response Preference Optimization with Augmented Ranking Dataset
by: Gwon, Hansle, et al.
Published: (2024)
by: Gwon, Hansle, et al.
Published: (2024)
NOTE: Notable generation Of patient Text summaries through Efficient approach based on direct preference optimization
by: Ahn, Imjin, et al.
Published: (2024)
by: Ahn, Imjin, et al.
Published: (2024)
Quality-Aware Translation Tagging in Multilingual RAG system
by: Moon, Hoyeon, et al.
Published: (2025)
by: Moon, Hoyeon, et al.
Published: (2025)
Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection
by: Song, Min Geun, et al.
Published: (2025)
by: Song, Min Geun, et al.
Published: (2025)
PHISH in MESH: Korean Adversarial Phonetic Substitution and Phonetic-Semantic Feature Integration Defense
by: Kim, Byungjun, et al.
Published: (2025)
by: Kim, Byungjun, et al.
Published: (2025)
Model Already Knows the Best Noise: Bayesian Active Noise Selection via Attention in Video Diffusion Model
by: Kim, Kwanyoung, et al.
Published: (2025)
by: Kim, Kwanyoung, et al.
Published: (2025)
StablePrompt: Automatic Prompt Tuning using Reinforcement Learning for Large Language Models
by: Kwon, Minchan, et al.
Published: (2024)
by: Kwon, Minchan, et al.
Published: (2024)
MultiFloodSynth: Multi-Annotated Flood Synthetic Dataset Generation
by: Kang, YoonJe, et al.
Published: (2025)
by: Kang, YoonJe, et al.
Published: (2025)
LIVE-GS: Online LiDAR-Inertial-Visual State Estimation and Globally Consistent Mapping with 3D Gaussian Splatting
by: Park, Jaeseok, et al.
Published: (2025)
by: Park, Jaeseok, et al.
Published: (2025)
Single-shot reconstruction of three-dimensional morphology of biological cells in digital holographic microscopy using a physics-driven neural network
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models
by: Park, Jungwon, et al.
Published: (2026)
by: Park, Jungwon, et al.
Published: (2026)
NexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
by: Kim, Hyuntak, et al.
Published: (2025)
by: Kim, Hyuntak, et al.
Published: (2025)
The ASIR Courage Model: A Phase-Dynamic Framework for Truth Transitions in Human and AI Systems
by: Kim, Hyo Jin
Published: (2026)
by: Kim, Hyo Jin
Published: (2026)
Investigating the Integrated Digital Interventions Delivered by a Therapeutic Companion Agent for Young Adults with Symptoms of Depression: A Proof-of-Concept Study
by: Yoo, Youngjae, et al.
Published: (2025)
by: Yoo, Youngjae, et al.
Published: (2025)
Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
Enhancing Exploration Efficiency using Uncertainty-Aware Information Prediction
by: Kim, Seunghwan, et al.
Published: (2024)
by: Kim, Seunghwan, et al.
Published: (2024)
A Highly Scalable TDMA for GPUs and Its Application to Flow Solver Optimization
by: Kim, Seungchan, et al.
Published: (2025)
by: Kim, Seungchan, et al.
Published: (2025)
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs
by: Kim, Sean, et al.
Published: (2025)
by: Kim, Sean, et al.
Published: (2025)
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models
by: Jung, Sunghee, et al.
Published: (2025)
by: Jung, Sunghee, et al.
Published: (2025)
Cylindrical Mechanical Projector for Omnidirectional Fringe Projection Profilometry
by: Choi, Mincheol, et al.
Published: (2026)
by: Choi, Mincheol, et al.
Published: (2026)
One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation
by: Jo, Sanghyun, et al.
Published: (2026)
by: Jo, Sanghyun, et al.
Published: (2026)
FRIDAY: Mitigating Unintentional Facial Identity in Deepfake Detectors Guided by Facial Recognizers
by: Kim, Younhun, et al.
Published: (2024)
by: Kim, Younhun, et al.
Published: (2024)
Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting Layers
by: Ko, Youmin, et al.
Published: (2025)
by: Ko, Youmin, et al.
Published: (2025)
IDF: Iterative Dynamic Filtering Networks for Generalizable Image Denoising
by: Kim, Dongjin, et al.
Published: (2025)
by: Kim, Dongjin, et al.
Published: (2025)
ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models
by: Yook, Hyun Jun, et al.
Published: (2025)
by: Yook, Hyun Jun, et al.
Published: (2025)
ToonAging: Face Re-Aging upon Artistic Portrait Style Transfer
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
SRHand: Super-Resolving Hand Images and 3D Shapes via View/Pose-aware Neural Image Representations and Explicit 3D Meshes
by: Kim, Minje, et al.
Published: (2025)
by: Kim, Minje, et al.
Published: (2025)
Arbitrary-Scale Image Generation and Upsampling using Latent Diffusion Model and Implicit Neural Decoder
by: Kim, Jinseok, et al.
Published: (2024)
by: Kim, Jinseok, et al.
Published: (2024)
LieHMR: Autoregressive Human Mesh Recovery with $SO(3)$ Diffusion
by: Kim, Donghwan, et al.
Published: (2025)
by: Kim, Donghwan, et al.
Published: (2025)
BiTT: Bi-directional Texture Reconstruction of Interacting Two Hands from a Single Image
by: Kim, Minje, et al.
Published: (2024)
by: Kim, Minje, et al.
Published: (2024)
Diffusion-Based sRGB Real Noise Generation via Prompt-Driven Noise Representation Learning
by: Ko, Jaekyun, et al.
Published: (2026)
by: Ko, Jaekyun, et al.
Published: (2026)
Designing LMS and Instructional Strategies for Integrating Generative-Conversational AI
by: Ra, Elias, et al.
Published: (2025)
by: Ra, Elias, et al.
Published: (2025)
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
by: Lee, Yoonjoo, et al.
Published: (2024)
by: Lee, Yoonjoo, et al.
Published: (2024)
Generalist Multi-Class Anomaly Detection via Distillation to Two Heterogeneous Student Networks
by: Park, Hangil, et al.
Published: (2025)
by: Park, Hangil, et al.
Published: (2025)
DiscoverLLM: From Executing Intents to Discovering Them
by: Kim, Tae Soo, et al.
Published: (2026)
by: Kim, Tae Soo, et al.
Published: (2026)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
by: Mu, Junjie, et al.
Published: (2025)
by: Mu, Junjie, et al.
Published: (2025)
Similar Items
-
Enhancing Clinical Efficiency through LLM: Discharge Note Generation for Cardiac Patients
by: Jung, HyoJe, et al.
Published: (2024) -
Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering
by: Kim, Byeolhee, et al.
Published: (2026) -
InMD-X: Large Language Models for Internal Medicine Doctors
by: Gwon, Hansle, et al.
Published: (2024) -
Multi-Response Preference Optimization with Augmented Ranking Dataset
by: Gwon, Hansle, et al.
Published: (2024) -
NOTE: Notable generation Of patient Text summaries through Efficient approach based on direct preference optimization
by: Ahn, Imjin, et al.
Published: (2024)