Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Yuting, Liu, Chengyuan, Feng, Yifeng, Wu, Yiquan, Wu, Chao, Wu, Fei, Kuang, Kun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios
von: Huang, Yuting, et al.
Veröffentlicht: (2025)
von: Huang, Yuting, et al.
Veröffentlicht: (2025)
P2S: Probabilistic Process Supervision for General-Domain Reasoning Question Answering
von: Zhong, Wenlin, et al.
Veröffentlicht: (2026)
von: Zhong, Wenlin, et al.
Veröffentlicht: (2026)
WisdomInterrogatory (LuWen): An Open-Source Legal Large Language Model Technical Report
von: Wu, Yiquan, et al.
Veröffentlicht: (2026)
von: Wu, Yiquan, et al.
Veröffentlicht: (2026)
PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs
von: Huang, Yuting, et al.
Veröffentlicht: (2026)
von: Huang, Yuting, et al.
Veröffentlicht: (2026)
Universal Legal Article Prediction via Tight Collaboration between Supervised Classification Model and LLM
von: Chi, Xiao, et al.
Veröffentlicht: (2025)
von: Chi, Xiao, et al.
Veröffentlicht: (2025)
From Graph to Word Bag: Introducing Domain Knowledge to Confusing Charge Prediction
von: Li, Ang, et al.
Veröffentlicht: (2024)
von: Li, Ang, et al.
Veröffentlicht: (2024)
SIGHT: Reinforcement Learning with Self-Evidence and Information-Gain Diverse Branching for Search Agent
von: Zhong, Wenlin, et al.
Veröffentlicht: (2026)
von: Zhong, Wenlin, et al.
Veröffentlicht: (2026)
Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
Evolving Knowledge Distillation with Large Language Models and Active Learning
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
Gold Panning in Vocabulary: An Adaptive Method for Vocabulary Expansion of Domain-Specific LLMs
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chiyu, et al.
Veröffentlicht: (2025)
More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
RexUniNLU: Recursive Method with Explicit Schema Instructor for Universal NLU
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Chengyuan, et al.
Veröffentlicht: (2024)
ClaimGen-CN: A Large-scale Chinese Dataset for Legal Claim Generation
von: Zhou, Siying, et al.
Veröffentlicht: (2025)
von: Zhou, Siying, et al.
Veröffentlicht: (2025)
Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers
von: Zhang, Tianhua, et al.
Veröffentlicht: (2024)
von: Zhang, Tianhua, et al.
Veröffentlicht: (2024)
You Know What I'm Saying: Jailbreak Attack via Implicit Reference
von: Wu, Tianyu, et al.
Veröffentlicht: (2024)
von: Wu, Tianyu, et al.
Veröffentlicht: (2024)
Intelligent Legal Assistant: An Interactive Clarification System for Legal Question Answering
von: Yao, Rujing, et al.
Veröffentlicht: (2025)
von: Yao, Rujing, et al.
Veröffentlicht: (2025)
Fine-tuning Large Language Models for Improving Factuality in Legal Question Answering
von: Hu, Yinghao, et al.
Veröffentlicht: (2025)
von: Hu, Yinghao, et al.
Veröffentlicht: (2025)
Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer
von: Kuang, Penghao, et al.
Veröffentlicht: (2026)
von: Kuang, Penghao, et al.
Veröffentlicht: (2026)
Automatic Slide Updating with User-Defined Dynamic Templates and Natural Language Instructions
von: Zhou, Kun, et al.
Veröffentlicht: (2026)
von: Zhou, Kun, et al.
Veröffentlicht: (2026)
Agentic Reinforcement Learning with Implicit Step Rewards
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2025)
Towards Stepwise Domain Knowledge-Driven Reasoning Optimization and Reflection Improvement
von: Liu, Chengyuan, et al.
Veröffentlicht: (2025)
von: Liu, Chengyuan, et al.
Veröffentlicht: (2025)
LLMs Encode Harmfulness and Refusal Separately
von: Zhao, Jiachen, et al.
Veröffentlicht: (2025)
von: Zhao, Jiachen, et al.
Veröffentlicht: (2025)
Target Span Detection for Implicit Harmful Content
von: Jafari, Nazanin, et al.
Veröffentlicht: (2024)
von: Jafari, Nazanin, et al.
Veröffentlicht: (2024)
Leveraging Print Debugging to Improve Code Generation in Large Language Models
von: Hu, Xueyu, et al.
Veröffentlicht: (2024)
von: Hu, Xueyu, et al.
Veröffentlicht: (2024)
Causal Agent based on Large Language Model
von: Han, Kairong, et al.
Veröffentlicht: (2024)
von: Han, Kairong, et al.
Veröffentlicht: (2024)
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond
von: Hu, Yinghao, et al.
Veröffentlicht: (2025)
von: Hu, Yinghao, et al.
Veröffentlicht: (2025)
Can Large Language Models Automatically Jailbreak GPT-4V?
von: Wu, Yuanwei, et al.
Veröffentlicht: (2024)
von: Wu, Yuanwei, et al.
Veröffentlicht: (2024)
Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful Knowledge
von: Lu, Weikai, et al.
Veröffentlicht: (2024)
von: Lu, Weikai, et al.
Veröffentlicht: (2024)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
LM-Combiner: A Contextual Rewriting Model for Chinese Grammatical Error Correction
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
DECOR: Improving Coherence in L2 English Writing with a Novel Benchmark for Incoherence Detection, Reasoning, and Rewriting
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
ModelGPT: Unleashing LLM's Capabilities for Tailored Model Generation
von: Tang, Zihao, et al.
Veröffentlicht: (2024)
von: Tang, Zihao, et al.
Veröffentlicht: (2024)
AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue
von: Park, Jihyung, et al.
Veröffentlicht: (2026)
von: Park, Jihyung, et al.
Veröffentlicht: (2026)
ImpScore: A Learnable Metric For Quantifying The Implicitness Level of Sentence
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
Tailoring Diagnostic Modeling to Individual Learners: Personalized Distractor Generation via MCTS-Guided Reasoning Reconstruction
von: Wu, Tao, et al.
Veröffentlicht: (2025)
von: Wu, Tao, et al.
Veröffentlicht: (2025)
Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting
von: Kang, Jingyi, et al.
Veröffentlicht: (2026)
von: Kang, Jingyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AppealCase: A Dataset and Benchmark for Civil Case Appeal Scenarios
von: Huang, Yuting, et al.
Veröffentlicht: (2025) -
P2S: Probabilistic Process Supervision for General-Domain Reasoning Question Answering
von: Zhong, Wenlin, et al.
Veröffentlicht: (2026) -
WisdomInterrogatory (LuWen): An Open-Source Legal Large Language Model Technical Report
von: Wu, Yiquan, et al.
Veröffentlicht: (2026) -
PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs
von: Huang, Yuting, et al.
Veröffentlicht: (2026) -
Universal Legal Article Prediction via Tight Collaboration between Supervised Classification Model and LLM
von: Chi, Xiao, et al.
Veröffentlicht: (2025)