Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Haoyu, Sun, Youran, Cai, Yunfeng, Zhu, Jun, Zhang, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What's the Magic Word? A Control Theory of LLM Prompting
von: Bhargava, Aman, et al.
Veröffentlicht: (2023)
von: Bhargava, Aman, et al.
Veröffentlicht: (2023)
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024)
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024)
Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
Text-Utilization for Encoder-dominated Speech Recognition Models
von: Zeyer, Albert, et al.
Veröffentlicht: (2026)
von: Zeyer, Albert, et al.
Veröffentlicht: (2026)
Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
Parameter-Efficient Fine-Tuning of LLMs with Mixture of Space Experts
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
von: Zhang, Buze, et al.
Veröffentlicht: (2026)
Hyperbolic Fine-Tuning for Large Language Models
von: Yang, Menglin, et al.
Veröffentlicht: (2024)
von: Yang, Menglin, et al.
Veröffentlicht: (2024)
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
von: Abhyankar, Nikhil, et al.
Veröffentlicht: (2025)
von: Abhyankar, Nikhil, et al.
Veröffentlicht: (2025)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
MAR: Efficient Large Language Models via Module-aware Architecture Refinement
von: Cai, Junhong, et al.
Veröffentlicht: (2026)
von: Cai, Junhong, et al.
Veröffentlicht: (2026)
BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models
von: Wang, Xinyuan, et al.
Veröffentlicht: (2024)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2024)
Recent Advances in Federated Learning Driven Large Language Models: A Survey on Architecture, Performance, and Security
von: Qu, Youyang, et al.
Veröffentlicht: (2024)
von: Qu, Youyang, et al.
Veröffentlicht: (2024)
Towards Universal Offline Black-Box Optimization via Learning Language Model Embeddings
von: Tan, Rong-Xi, et al.
Veröffentlicht: (2025)
von: Tan, Rong-Xi, et al.
Veröffentlicht: (2025)
EvoJail: Evolutionary Diverse Jailbreak Prompt Generation for Large Language Models
von: Tang, Rui, et al.
Veröffentlicht: (2026)
von: Tang, Rui, et al.
Veröffentlicht: (2026)
Evolutionary Computation in the Era of Large Language Model: Survey and Roadmap
von: Wu, Xingyu, et al.
Veröffentlicht: (2024)
von: Wu, Xingyu, et al.
Veröffentlicht: (2024)
Topic Modelling Black Box Optimization
von: Akramov, Roman, et al.
Veröffentlicht: (2025)
von: Akramov, Roman, et al.
Veröffentlicht: (2025)
Diffusion Language Models for Speech Recognition
von: Naveriani, Davyd, et al.
Veröffentlicht: (2026)
von: Naveriani, Davyd, et al.
Veröffentlicht: (2026)
Elastic Architecture Search for Efficient Language Models
von: Wang, Shang
Veröffentlicht: (2025)
von: Wang, Shang
Veröffentlicht: (2025)
EvoMerge: Neuroevolution for Large Language Models
von: Jiang, Yushu
Veröffentlicht: (2024)
von: Jiang, Yushu
Veröffentlicht: (2024)
Solve the Loop: Attractor Models for Language and Reasoning
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2026)
von: Fein-Ashley, Jacob, et al.
Veröffentlicht: (2026)
AgenticGEO: A Self-Evolving Agentic System for Generative Engine Optimization
von: Yuan, Jiaqi, et al.
Veröffentlicht: (2026)
von: Yuan, Jiaqi, et al.
Veröffentlicht: (2026)
H-Node Attack and Defense in Large Language Models
von: Yocam, Eric, et al.
Veröffentlicht: (2026)
von: Yocam, Eric, et al.
Veröffentlicht: (2026)
Fleet of Agents: Coordinated Problem Solving with Large Language Models
von: Klein, Lars, et al.
Veröffentlicht: (2024)
von: Klein, Lars, et al.
Veröffentlicht: (2024)
Large Language Models and Emergence: A Complex Systems Perspective
von: Krakauer, David C., et al.
Veröffentlicht: (2025)
von: Krakauer, David C., et al.
Veröffentlicht: (2025)
Assessing the Emergent Symbolic Reasoning Abilities of Llama Large Language Models
von: Petruzzellis, Flavio, et al.
Veröffentlicht: (2024)
von: Petruzzellis, Flavio, et al.
Veröffentlicht: (2024)
Evolutionary Multi-Objective Optimization of Large Language Model Prompts for Balancing Sentiments
von: Baumann, Jill, et al.
Veröffentlicht: (2024)
von: Baumann, Jill, et al.
Veröffentlicht: (2024)
Accelerating Training Speed of Tiny Recursive Models with Curriculum Guided Adaptive Recursion
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2025)
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2025)
LLM-SR: Scientific Equation Discovery via Programming with Large Language Models
von: Shojaee, Parshin, et al.
Veröffentlicht: (2024)
von: Shojaee, Parshin, et al.
Veröffentlicht: (2024)
When Large Language Models Meet Evolutionary Algorithms: Potential Enhancements and Challenges
von: Wang, Chao, et al.
Veröffentlicht: (2024)
von: Wang, Chao, et al.
Veröffentlicht: (2024)
Fine-Tuning and Evaluating Open-Source Large Language Models for the Army Domain
von: Ruiz, Daniel C., et al.
Veröffentlicht: (2024)
von: Ruiz, Daniel C., et al.
Veröffentlicht: (2024)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
von: Kaplan, Guy, et al.
Veröffentlicht: (2026)
von: Kaplan, Guy, et al.
Veröffentlicht: (2026)
W-PCA Based Gradient-Free Proxy for Efficient Search of Lightweight Language Models
von: Wang, Shang
Veröffentlicht: (2025)
von: Wang, Shang
Veröffentlicht: (2025)
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard
von: Topsakal, Oguzhan, et al.
Veröffentlicht: (2024)
von: Topsakal, Oguzhan, et al.
Veröffentlicht: (2024)
A Gauge Theory of Superposition: Toward a Sheaf-Theoretic Atlas of Neural Representations
von: Javidnia, Hossein
Veröffentlicht: (2026)
von: Javidnia, Hossein
Veröffentlicht: (2026)
Interlocking-free Selective Rationalization Through Genetic-based Learning
von: Ruggeri, Federico, et al.
Veröffentlicht: (2024)
von: Ruggeri, Federico, et al.
Veröffentlicht: (2024)
CAPO: Cost-Aware Prompt Optimization
von: Zehle, Tom, et al.
Veröffentlicht: (2025)
von: Zehle, Tom, et al.
Veröffentlicht: (2025)
Semantic Sections: An Atlas-Native Feature Ontology for Obstructed Representation Spaces
von: Javidnia, Hossein
Veröffentlicht: (2026)
von: Javidnia, Hossein
Veröffentlicht: (2026)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
von: Lee, Philip Heejun
Veröffentlicht: (2025)
von: Lee, Philip Heejun
Veröffentlicht: (2025)
EvoGPT-f: An Evolutionary GPT Framework for Benchmarking Formal Math Languages
von: Mercer, Johnathan
Veröffentlicht: (2024)
von: Mercer, Johnathan
Veröffentlicht: (2024)
Analysis of Error Sources in LLM-based Hypothesis Search for Few-Shot Rule Induction
von: Parab, Aishni, et al.
Veröffentlicht: (2025)
von: Parab, Aishni, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What's the Magic Word? A Control Theory of LLM Prompting
von: Bhargava, Aman, et al.
Veröffentlicht: (2023) -
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024) -
Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
von: Štefánik, Michal, et al.
Veröffentlicht: (2025) -
Text-Utilization for Encoder-dominated Speech Recognition Models
von: Zeyer, Albert, et al.
Veröffentlicht: (2026) -
Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits
von: Xing, Sixue, et al.
Veröffentlicht: (2026)