Black-Box Behavioral Distillation Breaks Safety Alignment in Medical LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Jahan, Sohely, Sun, Ruimin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
Bounded Behavioral Indistinguishability for Black-Box LLM Distillation
di: Hasan, Munawar
Pubblicazione: (2026)
di: Hasan, Munawar
Pubblicazione: (2026)
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
di: Rayhan, Naheed, et al.
Pubblicazione: (2026)
di: Rayhan, Naheed, et al.
Pubblicazione: (2026)
Boundary Point Jailbreaking of Black-Box LLMs
di: Davies, Xander, et al.
Pubblicazione: (2026)
di: Davies, Xander, et al.
Pubblicazione: (2026)
Turning Black Box into White Box: Dataset Distillation Leaks
di: Chen, Huajie, et al.
Pubblicazione: (2026)
di: Chen, Huajie, et al.
Pubblicazione: (2026)
Breaking Free: How to Hack Safety Guardrails in Black-Box Diffusion Models!
di: Kotyan, Shashank, et al.
Pubblicazione: (2024)
di: Kotyan, Shashank, et al.
Pubblicazione: (2024)
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
di: Fu, Yu, et al.
Pubblicazione: (2026)
di: Fu, Yu, et al.
Pubblicazione: (2026)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
di: Li, Changhao, et al.
Pubblicazione: (2024)
di: Li, Changhao, et al.
Pubblicazione: (2024)
Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment
di: Bakman, Yavuz, et al.
Pubblicazione: (2026)
di: Bakman, Yavuz, et al.
Pubblicazione: (2026)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
di: Xiao, Yuxin, et al.
Pubblicazione: (2025)
di: Xiao, Yuxin, et al.
Pubblicazione: (2025)
BridgePure: Limited Protection Leakage Can Break Black-Box Data Protection
di: Wang, Yihan, et al.
Pubblicazione: (2024)
di: Wang, Yihan, et al.
Pubblicazione: (2024)
Safety Filters for Black-Box Dynamical Systems by Learning Discriminating Hyperplanes
di: Lavanakul, Will, et al.
Pubblicazione: (2024)
di: Lavanakul, Will, et al.
Pubblicazione: (2024)
Enabling Fine-Grained Operating Points for Black-Box LLMs
di: Beyazit, Ege, et al.
Pubblicazione: (2025)
di: Beyazit, Ege, et al.
Pubblicazione: (2025)
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
di: Springer, Max, et al.
Pubblicazione: (2026)
di: Springer, Max, et al.
Pubblicazione: (2026)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
Bayesian Safety Validation for Failure Probability Estimation of Black-Box Systems
di: Moss, Robert J., et al.
Pubblicazione: (2023)
di: Moss, Robert J., et al.
Pubblicazione: (2023)
PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs
di: Chu, Jaewon, et al.
Pubblicazione: (2025)
di: Chu, Jaewon, et al.
Pubblicazione: (2025)
Multilingual Safety Alignment via Self-Distillation
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
Explaining the Behavior of Black-Box Prediction Algorithms with Causal Learning
di: Sani, Numair, et al.
Pubblicazione: (2020)
di: Sani, Numair, et al.
Pubblicazione: (2020)
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
di: Hong, Hanbin, et al.
Pubblicazione: (2023)
di: Hong, Hanbin, et al.
Pubblicazione: (2023)
Does Alignment Tuning Really Break LLMs' Internal Confidence?
di: Oh, Hongseok, et al.
Pubblicazione: (2024)
di: Oh, Hongseok, et al.
Pubblicazione: (2024)
GLiRA: Black-Box Membership Inference Attack via Knowledge Distillation
di: Galichin, Andrey V., et al.
Pubblicazione: (2024)
di: Galichin, Andrey V., et al.
Pubblicazione: (2024)
Black-Box Forgetting
di: Kuwana, Yusuke, et al.
Pubblicazione: (2024)
di: Kuwana, Yusuke, et al.
Pubblicazione: (2024)
Smoothing the Black-Box: Signed-Distance Supervision for Black-Box Model Copying
di: Jiménez, Rubén, et al.
Pubblicazione: (2026)
di: Jiménez, Rubén, et al.
Pubblicazione: (2026)
Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment
di: Vega, Jason, et al.
Pubblicazione: (2024)
di: Vega, Jason, et al.
Pubblicazione: (2024)
Beyond the Black Box: Interpretability of LLMs in Finance
di: Tatsat, Hariom, et al.
Pubblicazione: (2025)
di: Tatsat, Hariom, et al.
Pubblicazione: (2025)
Breaking the Reasoning Horizon in Entity Alignment Foundation Models
di: Cui, Yuanning, et al.
Pubblicazione: (2026)
di: Cui, Yuanning, et al.
Pubblicazione: (2026)
Aligning Logits Generatively for Principled Black-Box Knowledge Distillation
di: Ma, Jing, et al.
Pubblicazione: (2022)
di: Ma, Jing, et al.
Pubblicazione: (2022)
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
di: Salimian, Sina, et al.
Pubblicazione: (2025)
di: Salimian, Sina, et al.
Pubblicazione: (2025)
SafePassage: High-Fidelity Information Extraction with Black Box LLMs
di: Barrow, Joe, et al.
Pubblicazione: (2025)
di: Barrow, Joe, et al.
Pubblicazione: (2025)
ExecTune: Effective Steering of Black-Box LLMs with Guide Models
di: Lingam, Vijay, et al.
Pubblicazione: (2026)
di: Lingam, Vijay, et al.
Pubblicazione: (2026)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
di: Cho, Dongkyu Derek, et al.
Pubblicazione: (2025)
di: Cho, Dongkyu Derek, et al.
Pubblicazione: (2025)
Breaking the Black Box: Inherently Interpretable Physics-Constrained Machine Learning With Weighted Mixed-Effects for Imbalanced Seismic Data
di: Sreenath, Vemula, et al.
Pubblicazione: (2025)
di: Sreenath, Vemula, et al.
Pubblicazione: (2025)
Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
di: Agarwal, Krishiv, et al.
Pubblicazione: (2026)
di: Agarwal, Krishiv, et al.
Pubblicazione: (2026)
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
di: Zhang, Jiawei, et al.
Pubblicazione: (2025)
di: Zhang, Jiawei, et al.
Pubblicazione: (2025)
Black-Box Anomaly Attribution
di: Idé, Tsuyoshi, et al.
Pubblicazione: (2023)
di: Idé, Tsuyoshi, et al.
Pubblicazione: (2023)
FedAL: Black-Box Federated Knowledge Distillation Enabled by Adversarial Learning
di: Han, Pengchao, et al.
Pubblicazione: (2023)
di: Han, Pengchao, et al.
Pubblicazione: (2023)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
di: Abdulkadir, Perry
Pubblicazione: (2025)
di: Abdulkadir, Perry
Pubblicazione: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
di: Kroeger, Nicholas, et al.
Pubblicazione: (2023)
Hierarchical Support Vector State Partitioning for Distilling Black Box Reinforcement Learning Policies
di: Deproost, Senne, et al.
Pubblicazione: (2026)
di: Deproost, Senne, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
di: Nguyen, Tuan, et al.
Pubblicazione: (2025) -
Bounded Behavioral Indistinguishability for Black-Box LLM Distillation
di: Hasan, Munawar
Pubblicazione: (2026) -
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
di: Rayhan, Naheed, et al.
Pubblicazione: (2026) -
Boundary Point Jailbreaking of Black-Box LLMs
di: Davies, Xander, et al.
Pubblicazione: (2026) -
Turning Black Box into White Box: Dataset Distillation Leaks
di: Chen, Huajie, et al.
Pubblicazione: (2026)