Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
Fuente:
arXiv
Guardado en:
| Autores principales: | Yamabe, Shojiro, Sakuma, Jun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation
por: Yamabe, Shojiro, et al.
Publicado: (2024)
por: Yamabe, Shojiro, et al.
Publicado: (2024)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
por: Peng, Benji, et al.
Publicado: (2024)
por: Peng, Benji, et al.
Publicado: (2024)
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
por: Tran, Thien Q., et al.
Publicado: (2025)
por: Tran, Thien Q., et al.
Publicado: (2025)
MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language Models
por: Yamabe, Shojiro, et al.
Publicado: (2024)
por: Yamabe, Shojiro, et al.
Publicado: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
por: Aldahoul, Nouar, et al.
Publicado: (2025)
por: Aldahoul, Nouar, et al.
Publicado: (2025)
Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing Accuracy
por: Junhao, Wei, et al.
Publicado: (2025)
por: Junhao, Wei, et al.
Publicado: (2025)
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
por: Yamabe, Shojiro, et al.
Publicado: (2025)
por: Yamabe, Shojiro, et al.
Publicado: (2025)
Consensus Sampling for Safer Generative AI
por: Kalai, Adam Tauman, et al.
Publicado: (2025)
por: Kalai, Adam Tauman, et al.
Publicado: (2025)
DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
por: Majumder, Bodhisattwa Prasad, et al.
Publicado: (2024)
por: Majumder, Bodhisattwa Prasad, et al.
Publicado: (2024)
Deliberative Alignment: Reasoning Enables Safer Language Models
por: Guan, Melody Y., et al.
Publicado: (2024)
por: Guan, Melody Y., et al.
Publicado: (2024)
Certifiable Safe RLHF: Fixed-Penalty Constraint Optimization for Safer Language Models
por: Pandit, Kartik, et al.
Publicado: (2025)
por: Pandit, Kartik, et al.
Publicado: (2025)
On the MIA Vulnerability Gap Between Private GANs and Diffusion Models
por: Sebag, Ilana, et al.
Publicado: (2025)
por: Sebag, Ilana, et al.
Publicado: (2025)
AED: Automatic Discovery of Effective and Diverse Vulnerabilities for Autonomous Driving Policy with Large Language Models
por: Qiu, Le, et al.
Publicado: (2025)
por: Qiu, Le, et al.
Publicado: (2025)
SparseDM: Toward Sparse Efficient Diffusion Models
por: Wang, Kafeng, et al.
Publicado: (2024)
por: Wang, Kafeng, et al.
Publicado: (2024)
On Safer Reinforcement Learning for Sedation and Analgesia in Intensive Care
por: Romero-Hernandez, Joel, et al.
Publicado: (2026)
por: Romero-Hernandez, Joel, et al.
Publicado: (2026)
Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search
por: Ji, Junkai, et al.
Publicado: (2025)
por: Ji, Junkai, et al.
Publicado: (2025)
Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning
por: Hazra, Somnath, et al.
Publicado: (2025)
por: Hazra, Somnath, et al.
Publicado: (2025)
Network Level Evaluation of Hangup Susceptibility of HRGCs using Deep Learning and Sensing Techniques: A Goal Towards Safer Future
por: Chatterjee, Kaustav, et al.
Publicado: (2025)
por: Chatterjee, Kaustav, et al.
Publicado: (2025)
Laying Anchors: Semantically Priming Numerals in Language Modeling
por: Sharma, Mandar, et al.
Publicado: (2024)
por: Sharma, Mandar, et al.
Publicado: (2024)
Generating Counterfactual Trajectories with Latent Diffusion Models for Concept Discovery
por: Varshney, Payal, et al.
Publicado: (2024)
por: Varshney, Payal, et al.
Publicado: (2024)
Active Slice Discovery in Large Language Models
por: Zhang, Minhui, et al.
Publicado: (2025)
por: Zhang, Minhui, et al.
Publicado: (2025)
Nature Language Model: Deciphering the Language of Nature for Scientific Discovery
por: Xia, Yingce, et al.
Publicado: (2025)
por: Xia, Yingce, et al.
Publicado: (2025)
Free Draft-and-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models
por: Wu, Shutong, et al.
Publicado: (2025)
por: Wu, Shutong, et al.
Publicado: (2025)
Prime Convolutional Model: Breaking the Ground for Theoretical Explainability
por: Panelli, Francesco, et al.
Publicado: (2025)
por: Panelli, Francesco, et al.
Publicado: (2025)
Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness
por: Takahashi, Tsubasa, et al.
Publicado: (2025)
por: Takahashi, Tsubasa, et al.
Publicado: (2025)
Inference-Time Toxicity Mitigation in Protein Language Models
por: Burda, Manuel Fernández, et al.
Publicado: (2026)
por: Burda, Manuel Fernández, et al.
Publicado: (2026)
FlavorDiffusion: Predicting Food Pairings and Chemical Interactions Using Diffusion Models
por: Pyo, Seo Jun
Publicado: (2025)
por: Pyo, Seo Jun
Publicado: (2025)
Affective Priming Score: A Data-Driven Method to Detect Priming in Sequential Datasets
por: Maestro, Eduardo Gutierrez, et al.
Publicado: (2025)
por: Maestro, Eduardo Gutierrez, et al.
Publicado: (2025)
Statistically Significant Concept-based Explanation of Image Classifiers via Model Knockoffs
por: Xu, Kaiwen, et al.
Publicado: (2023)
por: Xu, Kaiwen, et al.
Publicado: (2023)
LookAhead Tuning: Safer Language Models via Partial Answer Previews
por: Liu, Kangwei, et al.
Publicado: (2025)
por: Liu, Kangwei, et al.
Publicado: (2025)
Large Language Models are Effective Priors for Causal Graph Discovery
por: Darvariu, Victor-Alexandru, et al.
Publicado: (2024)
por: Darvariu, Victor-Alexandru, et al.
Publicado: (2024)
Automated Attention Pattern Discovery at Scale in Large Language Models
por: Katzy, Jonathan, et al.
Publicado: (2026)
por: Katzy, Jonathan, et al.
Publicado: (2026)
Finetuning Large Language Models for Vulnerability Detection
por: Shestov, Alexey, et al.
Publicado: (2024)
por: Shestov, Alexey, et al.
Publicado: (2024)
Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?
por: Bengio, Yoshua, et al.
Publicado: (2025)
por: Bengio, Yoshua, et al.
Publicado: (2025)
Towards Mitigating Architecture Overfitting on Distilled Datasets
por: Zhong, Xuyang, et al.
Publicado: (2023)
por: Zhong, Xuyang, et al.
Publicado: (2023)
DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents
por: Zhao, Jiahao, et al.
Publicado: (2026)
por: Zhao, Jiahao, et al.
Publicado: (2026)
Mitigating Memorization In Language Models
por: Sakarvadia, Mansi, et al.
Publicado: (2024)
por: Sakarvadia, Mansi, et al.
Publicado: (2024)
Priming: Hybrid State Space Models From Pre-trained Transformers
por: Chattopadhyay, Aditya, et al.
Publicado: (2026)
por: Chattopadhyay, Aditya, et al.
Publicado: (2026)
TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models
por: Choo, Jinho, et al.
Publicado: (2026)
por: Choo, Jinho, et al.
Publicado: (2026)
LacMaterial: Large Language Models as Analogical Chemists for Materials Discovery
por: Guo, Hongyu
Publicado: (2025)
por: Guo, Hongyu
Publicado: (2025)
Ejemplares similares
-
Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation
por: Yamabe, Shojiro, et al.
Publicado: (2024) -
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
por: Peng, Benji, et al.
Publicado: (2024) -
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing
por: Tran, Thien Q., et al.
Publicado: (2025) -
MergePrint: Merge-Resistant Fingerprints for Robust Black-box Ownership Verification of Large Language Models
por: Yamabe, Shojiro, et al.
Publicado: (2024) -
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
por: Aldahoul, Nouar, et al.
Publicado: (2025)