Black-Box Behavioral Distillation Breaks Safety Alignment in Medical LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jahan, Sohely, Sun, Ruimin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Bounded Behavioral Indistinguishability for Black-Box LLM Distillation
von: Hasan, Munawar
Veröffentlicht: (2026)
von: Hasan, Munawar
Veröffentlicht: (2026)
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
von: Rayhan, Naheed, et al.
Veröffentlicht: (2026)
von: Rayhan, Naheed, et al.
Veröffentlicht: (2026)
Boundary Point Jailbreaking of Black-Box LLMs
von: Davies, Xander, et al.
Veröffentlicht: (2026)
von: Davies, Xander, et al.
Veröffentlicht: (2026)
Turning Black Box into White Box: Dataset Distillation Leaks
von: Chen, Huajie, et al.
Veröffentlicht: (2026)
von: Chen, Huajie, et al.
Veröffentlicht: (2026)
Breaking Free: How to Hack Safety Guardrails in Black-Box Diffusion Models!
von: Kotyan, Shashank, et al.
Veröffentlicht: (2024)
von: Kotyan, Shashank, et al.
Veröffentlicht: (2024)
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
von: Fu, Yu, et al.
Veröffentlicht: (2026)
von: Fu, Yu, et al.
Veröffentlicht: (2026)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
von: Li, Changhao, et al.
Veröffentlicht: (2024)
von: Li, Changhao, et al.
Veröffentlicht: (2024)
Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment
von: Bakman, Yavuz, et al.
Veröffentlicht: (2026)
von: Bakman, Yavuz, et al.
Veröffentlicht: (2026)
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
BridgePure: Limited Protection Leakage Can Break Black-Box Data Protection
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
Safety Filters for Black-Box Dynamical Systems by Learning Discriminating Hyperplanes
von: Lavanakul, Will, et al.
Veröffentlicht: (2024)
von: Lavanakul, Will, et al.
Veröffentlicht: (2024)
Enabling Fine-Grained Operating Points for Black-Box LLMs
von: Beyazit, Ege, et al.
Veröffentlicht: (2025)
von: Beyazit, Ege, et al.
Veröffentlicht: (2025)
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
von: Springer, Max, et al.
Veröffentlicht: (2026)
von: Springer, Max, et al.
Veröffentlicht: (2026)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
von: Chen, Xiwen, et al.
Veröffentlicht: (2026)
von: Chen, Xiwen, et al.
Veröffentlicht: (2026)
Bayesian Safety Validation for Failure Probability Estimation of Black-Box Systems
von: Moss, Robert J., et al.
Veröffentlicht: (2023)
von: Moss, Robert J., et al.
Veröffentlicht: (2023)
PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs
von: Chu, Jaewon, et al.
Veröffentlicht: (2025)
von: Chu, Jaewon, et al.
Veröffentlicht: (2025)
Multilingual Safety Alignment via Self-Distillation
von: Qin, Ruiyang, et al.
Veröffentlicht: (2026)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2026)
Explaining the Behavior of Black-Box Prediction Algorithms with Causal Learning
von: Sani, Numair, et al.
Veröffentlicht: (2020)
von: Sani, Numair, et al.
Veröffentlicht: (2020)
Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence
von: Hong, Hanbin, et al.
Veröffentlicht: (2023)
von: Hong, Hanbin, et al.
Veröffentlicht: (2023)
Does Alignment Tuning Really Break LLMs' Internal Confidence?
von: Oh, Hongseok, et al.
Veröffentlicht: (2024)
von: Oh, Hongseok, et al.
Veröffentlicht: (2024)
GLiRA: Black-Box Membership Inference Attack via Knowledge Distillation
von: Galichin, Andrey V., et al.
Veröffentlicht: (2024)
von: Galichin, Andrey V., et al.
Veröffentlicht: (2024)
Black-Box Forgetting
von: Kuwana, Yusuke, et al.
Veröffentlicht: (2024)
von: Kuwana, Yusuke, et al.
Veröffentlicht: (2024)
Smoothing the Black-Box: Signed-Distance Supervision for Black-Box Model Copying
von: Jiménez, Rubén, et al.
Veröffentlicht: (2026)
von: Jiménez, Rubén, et al.
Veröffentlicht: (2026)
Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment
von: Vega, Jason, et al.
Veröffentlicht: (2024)
von: Vega, Jason, et al.
Veröffentlicht: (2024)
Beyond the Black Box: Interpretability of LLMs in Finance
von: Tatsat, Hariom, et al.
Veröffentlicht: (2025)
von: Tatsat, Hariom, et al.
Veröffentlicht: (2025)
Breaking the Reasoning Horizon in Entity Alignment Foundation Models
von: Cui, Yuanning, et al.
Veröffentlicht: (2026)
von: Cui, Yuanning, et al.
Veröffentlicht: (2026)
Aligning Logits Generatively for Principled Black-Box Knowledge Distillation
von: Ma, Jing, et al.
Veröffentlicht: (2022)
von: Ma, Jing, et al.
Veröffentlicht: (2022)
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
von: Salimian, Sina, et al.
Veröffentlicht: (2025)
von: Salimian, Sina, et al.
Veröffentlicht: (2025)
SafePassage: High-Fidelity Information Extraction with Black Box LLMs
von: Barrow, Joe, et al.
Veröffentlicht: (2025)
von: Barrow, Joe, et al.
Veröffentlicht: (2025)
ExecTune: Effective Steering of Black-Box LLMs with Guide Models
von: Lingam, Vijay, et al.
Veröffentlicht: (2026)
von: Lingam, Vijay, et al.
Veröffentlicht: (2026)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
von: Cho, Dongkyu Derek, et al.
Veröffentlicht: (2025)
von: Cho, Dongkyu Derek, et al.
Veröffentlicht: (2025)
Breaking the Black Box: Inherently Interpretable Physics-Constrained Machine Learning With Weighted Mixed-Effects for Imbalanced Seismic Data
von: Sreenath, Vemula, et al.
Veröffentlicht: (2025)
von: Sreenath, Vemula, et al.
Veröffentlicht: (2025)
Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
von: Agarwal, Krishiv, et al.
Veröffentlicht: (2026)
von: Agarwal, Krishiv, et al.
Veröffentlicht: (2026)
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
Black-Box Anomaly Attribution
von: Idé, Tsuyoshi, et al.
Veröffentlicht: (2023)
von: Idé, Tsuyoshi, et al.
Veröffentlicht: (2023)
FedAL: Black-Box Federated Knowledge Distillation Enabled by Adversarial Learning
von: Han, Pengchao, et al.
Veröffentlicht: (2023)
von: Han, Pengchao, et al.
Veröffentlicht: (2023)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
von: Abdulkadir, Perry
Veröffentlicht: (2025)
von: Abdulkadir, Perry
Veröffentlicht: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
Hierarchical Support Vector State Partitioning for Distilling Black Box Reinforcement Learning Policies
von: Deproost, Senne, et al.
Veröffentlicht: (2026)
von: Deproost, Senne, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025) -
Bounded Behavioral Indistinguishability for Black-Box LLM Distillation
von: Hasan, Munawar
Veröffentlicht: (2026) -
Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models
von: Rayhan, Naheed, et al.
Veröffentlicht: (2026) -
Boundary Point Jailbreaking of Black-Box LLMs
von: Davies, Xander, et al.
Veröffentlicht: (2026) -
Turning Black Box into White Box: Dataset Distillation Leaks
von: Chen, Huajie, et al.
Veröffentlicht: (2026)