Boundary Point Jailbreaking of Black-Box LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Davies, Xander, Giglemiani, Giorgi, Lau, Edmund, Winsor, Eric, Irving, Geoffrey, Gal, Yarin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
von: Davies, Xander, et al.
Veröffentlicht: (2025)
von: Davies, Xander, et al.
Veröffentlicht: (2025)
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025)
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025)
Characterizing stable regions in the residual stream of LLMs
von: Janiak, Jett, et al.
Veröffentlicht: (2024)
von: Janiak, Jett, et al.
Veröffentlicht: (2024)
Do Multilingual LLMs Think In English?
von: Schut, Lisa, et al.
Veröffentlicht: (2025)
von: Schut, Lisa, et al.
Veröffentlicht: (2025)
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
von: Mehrotra, Anay, et al.
Veröffentlicht: (2023)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Jailbreaking Black Box Large Language Models in Twenty Queries
von: Chao, Patrick, et al.
Veröffentlicht: (2023)
von: Chao, Patrick, et al.
Veröffentlicht: (2023)
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
von: Souly, Alexandra, et al.
Veröffentlicht: (2025)
von: Souly, Alexandra, et al.
Veröffentlicht: (2025)
Enabling Fine-Grained Operating Points for Black-Box LLMs
von: Beyazit, Ege, et al.
Veröffentlicht: (2025)
von: Beyazit, Ege, et al.
Veröffentlicht: (2025)
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
von: Nikitin, Alexander, et al.
Veröffentlicht: (2024)
von: Nikitin, Alexander, et al.
Veröffentlicht: (2024)
SafePassage: High-Fidelity Information Extraction with Black Box LLMs
von: Barrow, Joe, et al.
Veröffentlicht: (2025)
von: Barrow, Joe, et al.
Veröffentlicht: (2025)
Simple Baselines are Competitive with Code Evolution
von: Gideoni, Yonatan, et al.
Veröffentlicht: (2026)
von: Gideoni, Yonatan, et al.
Veröffentlicht: (2026)
The Benefits and Risks of Transductive Approaches for AI Fairness
von: Razzak, Muhammed, et al.
Veröffentlicht: (2024)
von: Razzak, Muhammed, et al.
Veröffentlicht: (2024)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
von: Kossen, Jannik, et al.
Veröffentlicht: (2023)
von: Kossen, Jannik, et al.
Veröffentlicht: (2023)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
Temporal-Difference Variational Continual Learning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
Semantic Prototypes: Enhancing Transparency Without Black Boxes
von: Menis-Mastromichalakis, Orfeas, et al.
Veröffentlicht: (2024)
von: Menis-Mastromichalakis, Orfeas, et al.
Veröffentlicht: (2024)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
von: Kossen, Jannik, et al.
Veröffentlicht: (2024)
von: Kossen, Jannik, et al.
Veröffentlicht: (2024)
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
von: Li, Maximilian, et al.
Veröffentlicht: (2023)
von: Li, Maximilian, et al.
Veröffentlicht: (2023)
Training Transformers for KV Cache Compressibility
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
von: Gelberg, Yoav, et al.
Veröffentlicht: (2026)
Deep Bayesian Active Learning for Preference Modeling in Large Language Models
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
Iterative Deployment Improves Planning Skills in LLMs
von: Corrêa, Augusto B., et al.
Veröffentlicht: (2025)
von: Corrêa, Augusto B., et al.
Veröffentlicht: (2025)
PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs
von: Chu, Jaewon, et al.
Veröffentlicht: (2025)
von: Chu, Jaewon, et al.
Veröffentlicht: (2025)
Black-Box Behavioral Distillation Breaks Safety Alignment in Medical LLMs
von: Jahan, Sohely, et al.
Veröffentlicht: (2025)
von: Jahan, Sohely, et al.
Veröffentlicht: (2025)
Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions
von: Kim, Hazel, et al.
Veröffentlicht: (2024)
von: Kim, Hazel, et al.
Veröffentlicht: (2024)
Integrating White and Black Box Techniques for Interpretable Machine Learning
von: Vernon, Eric M., et al.
Veröffentlicht: (2024)
von: Vernon, Eric M., et al.
Veröffentlicht: (2024)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
von: Li, Changhao, et al.
Veröffentlicht: (2024)
von: Li, Changhao, et al.
Veröffentlicht: (2024)
Learning Antenna Pointing Correction in Operations: Efficient Calibration of a Black Box
von: Bergerhoff, Leif
Veröffentlicht: (2024)
von: Bergerhoff, Leif
Veröffentlicht: (2024)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
von: Sheng, Haotian, et al.
Veröffentlicht: (2026)
von: Sheng, Haotian, et al.
Veröffentlicht: (2026)
Boundary on the Table: Efficient Black-Box Decision-Based Attacks for Structured Data
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
von: Kazoom, Roie, et al.
Veröffentlicht: (2025)
Selective Safety Steering via Value-Filtered Decoding
von: Einbinder, Bat-Sheva, et al.
Veröffentlicht: (2026)
von: Einbinder, Bat-Sheva, et al.
Veröffentlicht: (2026)
Black-Box Forgetting
von: Kuwana, Yusuke, et al.
Veröffentlicht: (2024)
von: Kuwana, Yusuke, et al.
Veröffentlicht: (2024)
Smoothing the Black-Box: Signed-Distance Supervision for Black-Box Model Copying
von: Jiménez, Rubén, et al.
Veröffentlicht: (2026)
von: Jiménez, Rubén, et al.
Veröffentlicht: (2026)
Beyond the Black Box: Interpretability of LLMs in Finance
von: Tatsat, Hariom, et al.
Veröffentlicht: (2025)
von: Tatsat, Hariom, et al.
Veröffentlicht: (2025)
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
ExecTune: Effective Steering of Black-Box LLMs with Guide Models
von: Lingam, Vijay, et al.
Veröffentlicht: (2026)
von: Lingam, Vijay, et al.
Veröffentlicht: (2026)
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
von: Salimian, Sina, et al.
Veröffentlicht: (2025)
von: Salimian, Sina, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
von: Davies, Xander, et al.
Veröffentlicht: (2025) -
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
von: O'Brien, Kyle, et al.
Veröffentlicht: (2025) -
Characterizing stable regions in the residual stream of LLMs
von: Janiak, Jett, et al.
Veröffentlicht: (2024) -
Do Multilingual LLMs Think In English?
von: Schut, Lisa, et al.
Veröffentlicht: (2025) -
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
von: Giglemiani, Giorgi, et al.
Veröffentlicht: (2024)