xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Sunbowen, Ni, Shiwen, Wei, Chi, Li, Shuaimin, Fan, Liyang, Argha, Ahmadreza, Alinejad-Rokny, Hamid, Xu, Ruifeng, Gong, Yicheng, Yang, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
von: Li, Shuaimin, et al.
Veröffentlicht: (2025)
von: Li, Shuaimin, et al.
Veröffentlicht: (2025)
Interpretable graph-based models on multimodal biomedical data integration: A technical review and benchmarking
von: Sadeghi, Alireza, et al.
Veröffentlicht: (2025)
von: Sadeghi, Alireza, et al.
Veröffentlicht: (2025)
Lower Layers Matter: Alleviating Hallucination via Multi-Layer Fusion Contrastive Decoding with Truthfulness Refocused
von: Chen, Dingwei, et al.
Veröffentlicht: (2024)
von: Chen, Dingwei, et al.
Veröffentlicht: (2024)
CLinNET: An Interpretable and Uncertainty‐Aware Deep Learning Framework for Multi‐Modal Clinical Genomics
von: Ivan Bakhshayeshi, et al.
Veröffentlicht: (2026)
von: Ivan Bakhshayeshi, et al.
Veröffentlicht: (2026)
Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
von: Chen, Dingwei, et al.
Veröffentlicht: (2025)
von: Chen, Dingwei, et al.
Veröffentlicht: (2025)
ETAGE: Enhanced Test Time Adaptation with Integrated Entropy and Gradient Norms for Robust Model Performance
von: Shamsi, Afshar, et al.
Veröffentlicht: (2024)
von: Shamsi, Afshar, et al.
Veröffentlicht: (2024)
RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
SemanticST: Spatially Informed Semantic Graph Learning for Clustering, Integration, and Scalable Analysis of Spatial Transcriptomics
von: Zahedi, Roxana, et al.
Veröffentlicht: (2025)
von: Zahedi, Roxana, et al.
Veröffentlicht: (2025)
Small Language Model as Data Prospector for Large Language Model
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
AgentCourt: Simulating Court with Adversarial Evolvable Lawyer Agents
von: Chen, Guhong, et al.
Veröffentlicht: (2024)
von: Chen, Guhong, et al.
Veröffentlicht: (2024)
CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
von: Li, Siyi, et al.
Veröffentlicht: (2026)
von: Li, Siyi, et al.
Veröffentlicht: (2026)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
von: He, Zeqing, et al.
Veröffentlicht: (2024)
von: He, Zeqing, et al.
Veröffentlicht: (2024)
PersonaMath: Boosting Mathematical Reasoning via Persona-Driven Data Augmentation
von: Luo, Jing, et al.
Veröffentlicht: (2024)
von: Luo, Jing, et al.
Veröffentlicht: (2024)
JailbreakScope: Interpreting Jailbreak Mechanism through Representation and Circuit Analyses
von: He, Zeqing
Veröffentlicht: (2025)
von: He, Zeqing
Veröffentlicht: (2025)
Structuring Reasoning for Complex Rules Beyond Flat Representations
von: Yang, Zhihao, et al.
Veröffentlicht: (2025)
von: Yang, Zhihao, et al.
Veröffentlicht: (2025)
Counterfactual experience augmented off-policy reinforcement learning
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
Transcriptomic Models for Immunotherapy Response Prediction Show Limited Cross-cohort Generalisability
von: Liang, Yuheng, et al.
Veröffentlicht: (2026)
von: Liang, Yuheng, et al.
Veröffentlicht: (2026)
Quantification of Large Language Model Distillation
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
Probing the Difficulty Perception Mechanism of Large Language Models
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
CollectiveSFT: Scaling Large Language Models for Chinese Medical Benchmark with Collective Instructions in Healthcare
von: Zhu, Jingwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jingwei, et al.
Veröffentlicht: (2024)
How chromatin interactions shed light on interpreting non-coding genomic variants: opportunities and future direc-tions
von: Liang, Yuheng, et al.
Veröffentlicht: (2024)
von: Liang, Yuheng, et al.
Veröffentlicht: (2024)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
von: Ji, Haoxuan, et al.
Veröffentlicht: (2024)
von: Ji, Haoxuan, et al.
Veröffentlicht: (2024)
AutoPatent: A Multi-Agent Framework for Automatic Patent Generation
von: Wang, Qiyao, et al.
Veröffentlicht: (2024)
von: Wang, Qiyao, et al.
Veröffentlicht: (2024)
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning
von: Guo, Weiyang, et al.
Veröffentlicht: (2025)
von: Guo, Weiyang, et al.
Veröffentlicht: (2025)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
Jailbreaking to Jailbreak
von: Kritz, Jeremy, et al.
Veröffentlicht: (2025)
von: Kritz, Jeremy, et al.
Veröffentlicht: (2025)
TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts
von: Krauß, Torsten, et al.
Veröffentlicht: (2025)
von: Krauß, Torsten, et al.
Veröffentlicht: (2025)
Beyond Quantity: Trajectory Diversity Scaling for Code Agents
von: Chen, Guhong, et al.
Veröffentlicht: (2026)
von: Chen, Guhong, et al.
Veröffentlicht: (2026)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
von: Wei, Zhihua, et al.
Veröffentlicht: (2026)
von: Wei, Zhihua, et al.
Veröffentlicht: (2026)
Enhancing Monte Carlo Dropout Performance for Uncertainty Quantification
von: Asgharnezhad, Hamzeh, et al.
Veröffentlicht: (2025)
von: Asgharnezhad, Hamzeh, et al.
Veröffentlicht: (2025)
A Survey on Large Language Model Benchmarks
von: Ni, Shiwen, et al.
Veröffentlicht: (2025)
von: Ni, Shiwen, et al.
Veröffentlicht: (2025)
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Zhaoqi, et al.
Veröffentlicht: (2025)
Jailbreaking LLM-Controlled Robots
von: Robey, Alexander, et al.
Veröffentlicht: (2024)
von: Robey, Alexander, et al.
Veröffentlicht: (2024)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
von: Hossain, Ismail, et al.
Veröffentlicht: (2026)
FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis
von: Lin, Yuping, et al.
Veröffentlicht: (2024)
von: Lin, Yuping, et al.
Veröffentlicht: (2024)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
von: Zhou, Weikang, et al.
Veröffentlicht: (2024)
von: Zhou, Weikang, et al.
Veröffentlicht: (2024)
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
von: Wang, Qiyao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
von: Li, Shuaimin, et al.
Veröffentlicht: (2025) -
Interpretable graph-based models on multimodal biomedical data integration: A technical review and benchmarking
von: Sadeghi, Alireza, et al.
Veröffentlicht: (2025) -
Lower Layers Matter: Alleviating Hallucination via Multi-Layer Fusion Contrastive Decoding with Truthfulness Refocused
von: Chen, Dingwei, et al.
Veröffentlicht: (2024) -
CLinNET: An Interpretable and Uncertainty‐Aware Deep Learning Framework for Multi‐Modal Clinical Genomics
von: Ivan Bakhshayeshi, et al.
Veröffentlicht: (2026) -
Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
von: Chen, Dingwei, et al.
Veröffentlicht: (2025)