EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Xiaorui, Li, Fei, Mao, Xiaofeng, Zhang, Xin, Zheng, Li, Peng, Yuxiang, Teng, Chong, Ji, Donghong, Li, Zhuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis
by: Wu, Xiaorui, et al.
Published: (2025)
by: Wu, Xiaorui, et al.
Published: (2025)
Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering
by: Zheng, Li, et al.
Published: (2026)
by: Zheng, Li, et al.
Published: (2026)
DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinement
by: Lin, Shaoqing, et al.
Published: (2025)
by: Lin, Shaoqing, et al.
Published: (2025)
Revisiting Structured Sentiment Analysis as Latent Dependency Graph Parsing
by: Zhou, Chengjie, et al.
Published: (2024)
by: Zhou, Chengjie, et al.
Published: (2024)
CMNER: A Chinese Multimodal NER Dataset based on Social Media
by: Ji, Yuanze, et al.
Published: (2024)
by: Ji, Yuanze, et al.
Published: (2024)
Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge
by: Zheng, Li, et al.
Published: (2025)
by: Zheng, Li, et al.
Published: (2025)
Multi-Granular Multimodal Clue Fusion for Meme Understanding
by: Zheng, Li, et al.
Published: (2025)
by: Zheng, Li, et al.
Published: (2025)
PaSE: Prototype-aligned Calibration and Shapley-based Equilibrium for Multimodal Sentiment Analysis
by: He, Kang, et al.
Published: (2025)
by: He, Kang, et al.
Published: (2025)
Enhance-then-Balance Modality Collaboration for Robust Multimodal Sentiment Analysis
by: He, Kang, et al.
Published: (2026)
by: He, Kang, et al.
Published: (2026)
DALR: Dual-level Alignment Learning for Multimodal Sentence Representation Learning
by: He, Kang, et al.
Published: (2025)
by: He, Kang, et al.
Published: (2025)
Modeling Unified Semantic Discourse Structure for High-quality Headline Generation
by: Xu, Minghui, et al.
Published: (2024)
by: Xu, Minghui, et al.
Published: (2024)
Harvesting Events from Multiple Sources: Towards a Cross-Document Event Extraction Paradigm
by: Gao, Qiang, et al.
Published: (2024)
by: Gao, Qiang, et al.
Published: (2024)
Dynamic Emotion and Personality Profiling for Multimodal Deception Detection
by: Zheng, Li, et al.
Published: (2026)
by: Zheng, Li, et al.
Published: (2026)
Code-MIE: A Code-style Model for Multimodal Information Extraction with Scene Graph and Entity Attribute Knowledge Enhancement
by: Liu, Jiang, et al.
Published: (2026)
by: Liu, Jiang, et al.
Published: (2026)
Enhancing Cross-Document Event Coreference Resolution by Discourse Structure and Semantic Information
by: Gao, Qiang, et al.
Published: (2024)
by: Gao, Qiang, et al.
Published: (2024)
Zero-Shot Conversational Stance Detection: Dataset and Approaches
by: Ding, Yuzhe, et al.
Published: (2025)
by: Ding, Yuzhe, et al.
Published: (2025)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
by: An, Bang, et al.
Published: (2024)
by: An, Bang, et al.
Published: (2024)
Enhancing LLM Instruction Following: An Evaluation-Driven Multi-Agentic Workflow for Prompt Instructions Optimization
by: Purpura, Alberto, et al.
Published: (2026)
by: Purpura, Alberto, et al.
Published: (2026)
Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
by: Kang, Mintong, et al.
Published: (2025)
by: Kang, Mintong, et al.
Published: (2025)
Optimizing Instruction Synthesis: Effective Exploration of Evolutionary Space with Tree Search
by: Li, Chenglin, et al.
Published: (2024)
by: Li, Chenglin, et al.
Published: (2024)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
by: Young, Richard J., et al.
Published: (2026)
by: Young, Richard J., et al.
Published: (2026)
Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction
by: Yuan, Mengying, et al.
Published: (2025)
by: Yuan, Mengying, et al.
Published: (2025)
Consistency Assessment of CORDEX Multi‐Domain Simulations Over the Tibetan Plateau Using REMO
by: Ping Li, et al.
Published: (2025)
by: Ping Li, et al.
Published: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
Feature-Aware Malicious Output Detection and Mitigation
by: Dong, Weilong, et al.
Published: (2025)
by: Dong, Weilong, et al.
Published: (2025)
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
by: Jiang, Eric Hanchen, et al.
Published: (2025)
by: Jiang, Eric Hanchen, et al.
Published: (2025)
Smoke and Mirrors: Jailbreaking LLM-based Code Generation via Implicit Malicious Prompts
by: Ouyang, Sheng, et al.
Published: (2025)
by: Ouyang, Sheng, et al.
Published: (2025)
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Should LLM Safety Be More Than Refusing Harmful Instructions?
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation
by: Zhu, Runchuan, et al.
Published: (2025)
by: Zhu, Runchuan, et al.
Published: (2025)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
by: Wang, Peiran, et al.
Published: (2026)
by: Wang, Peiran, et al.
Published: (2026)
SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
by: Maskey, Utsav, et al.
Published: (2025)
by: Maskey, Utsav, et al.
Published: (2025)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
by: Zhang, Haonan, et al.
Published: (2025)
by: Zhang, Haonan, et al.
Published: (2025)
Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
by: Xiong, Junjie, et al.
Published: (2025)
by: Xiong, Junjie, et al.
Published: (2025)
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
by: Cheng, Ziheng, et al.
Published: (2025)
by: Cheng, Ziheng, et al.
Published: (2025)
Mitigating Malicious Attacks in Federated Learning via Confidence-aware Defense
by: Li, Qilei, et al.
Published: (2024)
by: Li, Qilei, et al.
Published: (2024)
Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive Refinement
by: Lu, Yuxiao, et al.
Published: (2026)
by: Lu, Yuxiao, et al.
Published: (2026)
Hear No Evil: Detecting Gradient Leakage by Malicious Servers in Federated Learning
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Similar Items
-
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis
by: Wu, Xiaorui, et al.
Published: (2025) -
Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering
by: Zheng, Li, et al.
Published: (2026) -
DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinement
by: Lin, Shaoqing, et al.
Published: (2025) -
Revisiting Structured Sentiment Analysis as Latent Dependency Graph Parsing
by: Zhou, Chengjie, et al.
Published: (2024) -
CMNER: A Chinese Multimodal NER Dataset based on Social Media
by: Ji, Yuanze, et al.
Published: (2024)