Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
Fuente:
arXiv
Saved in:
| Main Authors: | Tsmindashvili, Tatia, Kolkhidashvili, Ana, Kurtskhalia, Dachi, Maghlakelidze, Nino, Mekvabishvili, Elene, Dentoshvili, Guram, Shamilov, Orkhan, Gachechiladze, Zaal, Saporta, Steven, Choladze, David Dachi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GREEN CONSUMER BEHAVIOR: EVIDENCE FROM THE BRAZIL – URUGUAY BORDER REGION
by: Juliana Dachi Vieira
Published: (2019)
by: Juliana Dachi Vieira
Published: (2019)
Accelerated Cloud for Artificial Intelligence (ACAI)
by: Chen, Dachi, et al.
Published: (2024)
by: Chen, Dachi, et al.
Published: (2024)
Merging Improves Self-Critique Against Jailbreak Attacks
by: Gallego, Victor
Published: (2024)
by: Gallego, Victor
Published: (2024)
Decorating Pd–Au Nanodots Around Porous In2O3 Nanocubes for Tolerant H2 Sensing Against Switching Response and H2S Poisoning
by: Xinhua Zhao, et al.
Published: (2024)
by: Xinhua Zhao, et al.
Published: (2024)
Cyanogel‐Transformed Porous Palladium and Iron Framework Intermixed with rGO for Wearable Hydrogen Sensing
by: Xinhua Zhao, et al.
Published: (2024)
by: Xinhua Zhao, et al.
Published: (2024)
Resonances in a Dirichlet quantum waveguide coupled to a cavity
by: Kondej, Sylwia, et al.
Published: (2026)
by: Kondej, Sylwia, et al.
Published: (2026)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
Words of Power
by: Mir-Kasimov, Orkhan
Published: (2025)
by: Mir-Kasimov, Orkhan
Published: (2025)
My journey with Kamala Kempadoo: Reflections on the anti‐trafficking movement and sex work organizing
by: Elene Lam
Published: (2025)
by: Elene Lam
Published: (2025)
Cooperation After the Algorithm: Designing Human-AI Coexistence Beyond the Illusion of Collaboration
by: Codreanu, Tatia
Published: (2026)
by: Codreanu, Tatia
Published: (2026)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)
by: Feng, Yingchaojie, et al.
Published: (2024)
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
by: Wei, Zhipeng, et al.
Published: (2024)
by: Wei, Zhipeng, et al.
Published: (2024)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
by: Jaiswal, Piyush, et al.
Published: (2026)
by: Jaiswal, Piyush, et al.
Published: (2026)
Voice Jailbreak Attacks Against GPT-4o
by: Shen, Xinyue, et al.
Published: (2024)
by: Shen, Xinyue, et al.
Published: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
by: Li, Yanzeng, et al.
Published: (2025)
by: Li, Yanzeng, et al.
Published: (2025)
Einsatz generativer KI-Systeme im Unterricht
by: Jalilov, Orkhan, et al.
Published: (2025)
by: Jalilov, Orkhan, et al.
Published: (2025)
Hard Rock Drilling for Super-hot Enhanced Geothermal System Development: Literature Review and Techno-Economic Analysis
by: Khankishiyev, Orkhan, et al.
Published: (2024)
by: Khankishiyev, Orkhan, et al.
Published: (2024)
Distribution Transformer Winding Fault Detection Based on Hybrid Wavelet‐CNN
by: Orkhan Afandi, et al.
Published: (2025)
by: Orkhan Afandi, et al.
Published: (2025)
Evaluation in EEG Emotion Recognition: State-of-the-Art Review and Unified Framework
by: Kukhilava, Natia, et al.
Published: (2025)
by: Kukhilava, Natia, et al.
Published: (2025)
AlignTree: Efficient Defense Against LLM Jailbreak Attacks
by: Goren, Gil, et al.
Published: (2025)
by: Goren, Gil, et al.
Published: (2025)
Seleção e preferência de microhábitats por larvas de peixes migradores.
by: Taguti, Tátia Leika
Published: (2011)
by: Taguti, Tátia Leika
Published: (2011)
Early development of two tropical fishes (Perciformes: Sciaenidae) from the Pantanal of Mato Grosso, Brazil
by: Tátia Leika Taguti
Published: (2015)
by: Tátia Leika Taguti
Published: (2015)
Impact of female genital mutilation laws, policies, and professional codes of conduct on healthcare workers' knowledge, attitudes, skills, and quality of care: A mixed‐method review
by: Ibitola Asaolu, et al.
Published: (2026)
by: Ibitola Asaolu, et al.
Published: (2026)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
by: Nikolić, Kristina, et al.
Published: (2025)
by: Nikolić, Kristina, et al.
Published: (2025)
Modified light-cylinder and centrifugal acceleration in Schwarzschild geometry
by: Kurtskhalia, Nikoloz, et al.
Published: (2025)
by: Kurtskhalia, Nikoloz, et al.
Published: (2025)
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
by: Ran, Delong, et al.
Published: (2024)
by: Ran, Delong, et al.
Published: (2024)
Correspondence to “A preliminary study of collaborative group intervention with recovered peer supporters for eating disorders: Analyses including comparisons between in‐person and online sessions”
by: Nirjal Thapa, et al.
Published: (2024)
by: Nirjal Thapa, et al.
Published: (2024)
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
by: Robey, Alexander, et al.
Published: (2023)
by: Robey, Alexander, et al.
Published: (2023)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
by: Liu, Xiaoqun, et al.
Published: (2024)
by: Liu, Xiaoqun, et al.
Published: (2024)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
by: Zhou, Andy, et al.
Published: (2024)
by: Zhou, Andy, et al.
Published: (2024)
Low-Effort Jailbreak Attacks Against Text-to-Image Safety Filters
by: Mustafa, Ahmed B, et al.
Published: (2026)
by: Mustafa, Ahmed B, et al.
Published: (2026)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
by: Yin, Ziyi, et al.
Published: (2025)
by: Yin, Ziyi, et al.
Published: (2025)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
by: Yi, Sibo, et al.
Published: (2024)
by: Yi, Sibo, et al.
Published: (2024)
DNA Molecular Glue Assisted Bacterial Conjugative Transfer
by: Liqing Qi, et al.
Published: (2024)
by: Liqing Qi, et al.
Published: (2024)
High Schmidt number concentration in quantum bound entangled states
by: Krebs, Robin, et al.
Published: (2024)
by: Krebs, Robin, et al.
Published: (2024)
Scaling Bound Entanglement through Local Extensions
by: Krebs, Robin, et al.
Published: (2025)
by: Krebs, Robin, et al.
Published: (2025)
SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
by: Li, Xiangman, et al.
Published: (2025)
by: Li, Xiangman, et al.
Published: (2025)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
by: Zhang, Zhexin, et al.
Published: (2023)
by: Zhang, Zhexin, et al.
Published: (2023)
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Similar Items
-
GREEN CONSUMER BEHAVIOR: EVIDENCE FROM THE BRAZIL – URUGUAY BORDER REGION
by: Juliana Dachi Vieira
Published: (2019) -
Accelerated Cloud for Artificial Intelligence (ACAI)
by: Chen, Dachi, et al.
Published: (2024) -
Merging Improves Self-Critique Against Jailbreak Attacks
by: Gallego, Victor
Published: (2024) -
Decorating Pd–Au Nanodots Around Porous In2O3 Nanocubes for Tolerant H2 Sensing Against Switching Response and H2S Poisoning
by: Xinhua Zhao, et al.
Published: (2024) -
Cyanogel‐Transformed Porous Palladium and Iron Framework Intermixed with rGO for Wearable Hydrogen Sensing
by: Xinhua Zhao, et al.
Published: (2024)