Feature-Aware Malicious Output Detection and Mitigation
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Weilong, Li, Peiguang, Tian, Yu, Zeng, Xinyi, Li, Fengdi, Wang, Sirui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs Know What They Need: Leveraging a Missing Information Guided Framework to Empower Retrieval-Augmented Generation
by: Wang, Keheng, et al.
Published: (2024)
by: Wang, Keheng, et al.
Published: (2024)
Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
by: Tian, Changyuan, et al.
Published: (2025)
by: Tian, Changyuan, et al.
Published: (2025)
Latent Distribution Decoupling: A Probabilistic Framework for Uncertainty-Aware Multimodal Emotion Recognition
by: Huang, Jingwang, et al.
Published: (2025)
by: Huang, Jingwang, et al.
Published: (2025)
IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware Neurons
by: Shi, Dan, et al.
Published: (2024)
by: Shi, Dan, et al.
Published: (2024)
Localizing Malicious Outputs from CodeLLM
by: Borana, Mayukh, et al.
Published: (2025)
by: Borana, Mayukh, et al.
Published: (2025)
Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems
by: Yang, Ziyuan, et al.
Published: (2026)
by: Yang, Ziyuan, et al.
Published: (2026)
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
by: Wang, Huandong, et al.
Published: (2025)
by: Wang, Huandong, et al.
Published: (2025)
Uncertainty Unveiled: Can Exposure to More In-context Examples Mitigate Uncertainty for Large Language Models?
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
by: Long, Do Xuan, et al.
Published: (2024)
by: Long, Do Xuan, et al.
Published: (2024)
Truth-Aware Context Selection: Mitigating Hallucinations of Large Language Models Being Misled by Untruthful Contexts
by: Yu, Tian, et al.
Published: (2024)
by: Yu, Tian, et al.
Published: (2024)
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
by: Shen, Xinjie, et al.
Published: (2026)
by: Shen, Xinjie, et al.
Published: (2026)
"Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
Semantic and Contextual Modeling for Malicious Comment Detection with BERT-BiLSTM
by: Fang, Zhou, et al.
Published: (2025)
by: Fang, Zhou, et al.
Published: (2025)
On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMs
by: Wan, Herun, et al.
Published: (2024)
by: Wan, Herun, et al.
Published: (2024)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
by: Gur-Arieh, Yoav, et al.
Published: (2025)
by: Gur-Arieh, Yoav, et al.
Published: (2025)
Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
Mitigating Biases of Large Language Models in Stance Detection with Counterfactual Augmented Calibration
by: Li, Ang, et al.
Published: (2024)
by: Li, Ang, et al.
Published: (2024)
Disclosure and Mitigation of Gender Bias in LLMs
by: Dong, Xiangjue, et al.
Published: (2024)
by: Dong, Xiangjue, et al.
Published: (2024)
Reformulation is All You Need: Addressing Malicious Text Features in DNNs
by: Jiang, Yi, et al.
Published: (2025)
by: Jiang, Yi, et al.
Published: (2025)
LLM-Driven Reasoning for Constraint-Aware Feature Selection in Industrial Systems
by: Zhou, Yuhang, et al.
Published: (2026)
by: Zhou, Yuhang, et al.
Published: (2026)
DiffER: Diffusion Entity-Relation Modeling for Reversal Curse in Diffusion Large Language Models
by: He, Shaokai, et al.
Published: (2026)
by: He, Shaokai, et al.
Published: (2026)
To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing
by: Cheng, Wei, et al.
Published: (2026)
by: Cheng, Wei, et al.
Published: (2026)
Compile to Compress: Boosting Formal Theorem Provers by Compiler Outputs
by: Li, Guchan, et al.
Published: (2026)
by: Li, Guchan, et al.
Published: (2026)
Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
by: Zeng, Xinyi, et al.
Published: (2024)
by: Zeng, Xinyi, et al.
Published: (2024)
Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection
by: Wan, Herun, et al.
Published: (2025)
by: Wan, Herun, et al.
Published: (2025)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
by: Watts, Ishaan, et al.
Published: (2026)
by: Watts, Ishaan, et al.
Published: (2026)
Exploring Multilingual Concepts of Human Value in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages?
by: Xu, Shaoyang, et al.
Published: (2024)
by: Xu, Shaoyang, et al.
Published: (2024)
ConTrans: Weak-to-Strong Alignment Engineering via Concept Transplantation
by: Dong, Weilong, et al.
Published: (2024)
by: Dong, Weilong, et al.
Published: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
by: Chen, Junjie, et al.
Published: (2026)
by: Chen, Junjie, et al.
Published: (2026)
DRS: Deep Question Reformulation With Structured Output
by: Li, Zhecheng, et al.
Published: (2024)
by: Li, Zhecheng, et al.
Published: (2024)
Entity-Aware Self-Attention and Contextualized GCN for Enhanced Relation Extraction in Long Sentences
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
by: Zhang, Zhibo, et al.
Published: (2024)
by: Zhang, Zhibo, et al.
Published: (2024)
Group-Adaptive Adversarial Learning for Robust Fake News Detection Against Malicious Comments
by: Tong, Zhao, et al.
Published: (2025)
by: Tong, Zhao, et al.
Published: (2025)
Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation
by: Fadeeva, Ekaterina, et al.
Published: (2025)
by: Fadeeva, Ekaterina, et al.
Published: (2025)
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance
by: Fan, Yuchun, et al.
Published: (2026)
by: Fan, Yuchun, et al.
Published: (2026)
Developing a Reliable, Fast, General-Purpose Hallucination Detection and Mitigation Service
by: Wang, Song, et al.
Published: (2024)
by: Wang, Song, et al.
Published: (2024)
Me-Agent: A Personalized Mobile Agent with Two-Level User Habit Learning for Enhanced Interaction
by: Wang, Shuoxin, et al.
Published: (2026)
by: Wang, Shuoxin, et al.
Published: (2026)
RAVE: Retrieval and Scoring Aware Verifiable Claim Detection
by: Li, Yufeng, et al.
Published: (2025)
by: Li, Yufeng, et al.
Published: (2025)
Similar Items
-
LLMs Know What They Need: Leveraging a Missing Information Guided Framework to Empower Retrieval-Augmented Generation
by: Wang, Keheng, et al.
Published: (2024) -
Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
by: Tian, Changyuan, et al.
Published: (2025) -
Latent Distribution Decoupling: A Probabilistic Framework for Uncertainty-Aware Multimodal Emotion Recognition
by: Huang, Jingwang, et al.
Published: (2025) -
IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware Neurons
by: Shi, Dan, et al.
Published: (2024) -
Localizing Malicious Outputs from CodeLLM
by: Borana, Mayukh, et al.
Published: (2025)