ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
Fuente:
arXiv
Salvato in:
| Autori principali: | Kang, Hankun, Miao, Xin, Chen, Jianhao, Wen, Jintao, Xu, Mayi, Zhang, Weiyu, Lu, Wenpeng, Qian, Tieyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Toxicity Detection towards Adaptability to Changing Perturbations
di: Kang, Hankun, et al.
Pubblicazione: (2024)
di: Kang, Hankun, et al.
Pubblicazione: (2024)
Aligning VLM Assistants with Personalized Situated Cognition
di: Li, Yongqi, et al.
Pubblicazione: (2025)
di: Li, Yongqi, et al.
Pubblicazione: (2025)
Reasoning based on symbolic and parametric knowledge bases: a survey
di: Xu, Mayi, et al.
Pubblicazione: (2025)
di: Xu, Mayi, et al.
Pubblicazione: (2025)
Prompting Large Language Models for Counterfactual Generation: An Empirical Study
di: Li, Yongqi, et al.
Pubblicazione: (2023)
di: Li, Yongqi, et al.
Pubblicazione: (2023)
A Survey on Training-free Alignment of Large Language Models
di: Pan, Birong, et al.
Pubblicazione: (2025)
di: Pan, Birong, et al.
Pubblicazione: (2025)
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
di: Li, Xiaohu, et al.
Pubblicazione: (2025)
di: Li, Xiaohu, et al.
Pubblicazione: (2025)
Enhancing Relation Extraction via Supervised Rationale Verification and Feedback
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
Can a Small Model Learn to Look Before It Leaps? Dynamic Learning and Proactive Correction for Hallucination Detection
di: Bao, Zepeng, et al.
Pubblicazione: (2025)
di: Bao, Zepeng, et al.
Pubblicazione: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
di: Pan, Birong, et al.
Pubblicazione: (2025)
di: Pan, Birong, et al.
Pubblicazione: (2025)
Privacy-protected Retrieval-Augmented Generation for Knowledge Graph Question Answering
di: Ning, Yunfeng, et al.
Pubblicazione: (2025)
di: Ning, Yunfeng, et al.
Pubblicazione: (2025)
RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
di: Chen, Jianhao, et al.
Pubblicazione: (2025)
di: Chen, Jianhao, et al.
Pubblicazione: (2025)
Beyond Static Snapshots: Dynamic Modeling and Forecasting of Group-Level Value Evolution with Large Language Models
di: Pi, Qiankun, et al.
Pubblicazione: (2026)
di: Pi, Qiankun, et al.
Pubblicazione: (2026)
Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
di: Liu, Jiacheng, et al.
Pubblicazione: (2025)
di: Liu, Jiacheng, et al.
Pubblicazione: (2025)
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
di: Chen, Jianhao, et al.
Pubblicazione: (2026)
di: Chen, Jianhao, et al.
Pubblicazione: (2026)
AdaCultureSafe: Adaptive Cultural Safety Grounded by Cultural Knowledge in Large Language Models
di: Kang, Hankun, et al.
Pubblicazione: (2026)
di: Kang, Hankun, et al.
Pubblicazione: (2026)
KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing
di: Xu, Zhenhua, et al.
Pubblicazione: (2026)
di: Xu, Zhenhua, et al.
Pubblicazione: (2026)
Conti-Fuse: A Novel Continuous Decomposition-based Fusion Framework for Infrared and Visible Images
di: Li, Hui, et al.
Pubblicazione: (2024)
di: Li, Hui, et al.
Pubblicazione: (2024)
Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
di: Xu, Wenpeng
Pubblicazione: (2026)
di: Xu, Wenpeng
Pubblicazione: (2026)
ContiFormer: Continuous-Time Transformer for Irregular Time Series Modeling
di: Chen, Yuqi, et al.
Pubblicazione: (2024)
di: Chen, Yuqi, et al.
Pubblicazione: (2024)
Knowledge Graph Tokenization for Behavior-Aware Generative Next POI Recommendation
di: Sun, Ke, et al.
Pubblicazione: (2025)
di: Sun, Ke, et al.
Pubblicazione: (2025)
In-Application Defense Against Evasive Web Scans through Behavioral Analysis
di: Ousat, Behzad, et al.
Pubblicazione: (2024)
di: Ousat, Behzad, et al.
Pubblicazione: (2024)
Michelangelo / Giulio Carlo Argan, Bruno Conti
di: Argan, Giulio C
Pubblicazione: (1987)
di: Argan, Giulio C
Pubblicazione: (1987)
The Protective Role of Ambient Ultraviolet Radiation Against Dementia: An Ecological Analysis of Global Data
di: Wenpeng You
Pubblicazione: (2025)
di: Wenpeng You
Pubblicazione: (2025)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
di: Li, Siyuan, et al.
Pubblicazione: (2026)
di: Li, Siyuan, et al.
Pubblicazione: (2026)
Guardians of the Network: An Ensemble Learning Framework With Adversarial Alignment for Evasive Cyber Threat Detection
di: Khandakar Md Shafin, et al.
Pubblicazione: (2025)
di: Khandakar Md Shafin, et al.
Pubblicazione: (2025)
Local development and competitiveness / Sergio Conti ; Paolo Giaccaria
di: Conti, Sergio
di: Conti, Sergio
EvoGuard: An Extensible Agentic RL-based Framework for Practical and Evolving AI-Generated Image Detection
di: Zhu, Chenyang, et al.
Pubblicazione: (2026)
di: Zhu, Chenyang, et al.
Pubblicazione: (2026)
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents
di: Kim, San, et al.
Pubblicazione: (2024)
di: Kim, San, et al.
Pubblicazione: (2024)
Adversarial Pre-Padding: Generating Evasive Network Traffic Against Transformer-Based Classifiers
di: Jing, Quanliang, et al.
Pubblicazione: (2025)
di: Jing, Quanliang, et al.
Pubblicazione: (2025)
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
di: Xu, Ancheng, et al.
Pubblicazione: (2025)
di: Xu, Ancheng, et al.
Pubblicazione: (2025)
POSTER: A Multi-Signal Model for Detecting Evasive Smishing
di: Hosseinpour, Shaghayegh, et al.
Pubblicazione: (2025)
di: Hosseinpour, Shaghayegh, et al.
Pubblicazione: (2025)
Variety Evasive Subspace Families
di: Guo, Zeyu
Pubblicazione: (2021)
di: Guo, Zeyu
Pubblicazione: (2021)
A Unified and Scalable Algorithm Framework of User-Defined Temporal $(k,\mathcal{X})$-Core Query
di: Zhong, Ming, et al.
Pubblicazione: (2023)
di: Zhong, Ming, et al.
Pubblicazione: (2023)
Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions
di: Li, Yongqi, et al.
Pubblicazione: (2026)
di: Li, Yongqi, et al.
Pubblicazione: (2026)
Megatron: Evasive Clean-Label Backdoor Attacks against Vision Transformer
di: Gong, Xueluan, et al.
Pubblicazione: (2024)
di: Gong, Xueluan, et al.
Pubblicazione: (2024)
Evasive Random Walks and the Clairvoyant Demon
di: Abrams, Aaron, et al.
Pubblicazione: (2025)
di: Abrams, Aaron, et al.
Pubblicazione: (2025)
SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation
di: Su, Guangzhi, et al.
Pubblicazione: (2025)
di: Su, Guangzhi, et al.
Pubblicazione: (2025)
RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection
di: Zhu, He, et al.
Pubblicazione: (2026)
di: Zhu, He, et al.
Pubblicazione: (2026)
Guarding Terrains with Guards on a Line
di: Kang, Byeonguk, et al.
Pubblicazione: (2025)
di: Kang, Byeonguk, et al.
Pubblicazione: (2025)
Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual Learning
di: Qiu, Haomiao, et al.
Pubblicazione: (2025)
di: Qiu, Haomiao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Toxicity Detection towards Adaptability to Changing Perturbations
di: Kang, Hankun, et al.
Pubblicazione: (2024) -
Aligning VLM Assistants with Personalized Situated Cognition
di: Li, Yongqi, et al.
Pubblicazione: (2025) -
Reasoning based on symbolic and parametric knowledge bases: a survey
di: Xu, Mayi, et al.
Pubblicazione: (2025) -
Prompting Large Language Models for Counterfactual Generation: An Empirical Study
di: Li, Yongqi, et al.
Pubblicazione: (2023) -
A Survey on Training-free Alignment of Large Language Models
di: Pan, Birong, et al.
Pubblicazione: (2025)