Efficient LLM Moderation with Multi-Layer Latent Prototypes
Fuente:
arXiv
Salvato in:
| Autori principali: | Chrabąszcz, Maciej, Szatkowski, Filip, Wójcik, Bartosz, Dubiński, Jan, Trzciński, Tomasz, Cygert, Sebastian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2026)
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2026)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
di: Pawlak, Stanisław, et al.
Pubblicazione: (2025)
di: Pawlak, Stanisław, et al.
Pubblicazione: (2025)
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
di: Dubiński, Jan, et al.
Pubblicazione: (2026)
di: Dubiński, Jan, et al.
Pubblicazione: (2026)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
di: Liang, Buyun, et al.
Pubblicazione: (2026)
di: Liang, Buyun, et al.
Pubblicazione: (2026)
FLAME: Flexible LLM-Assisted Moderation Engine
di: Bakulin, Ivan, et al.
Pubblicazione: (2025)
di: Bakulin, Ivan, et al.
Pubblicazione: (2025)
Addressing The Devastating Effects Of Single-Task Data Poisoning In Exemplar-Free Continual Learning
di: Pawlak, Stanisław, et al.
Pubblicazione: (2025)
di: Pawlak, Stanisław, et al.
Pubblicazione: (2025)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
di: Huang, Tiansheng, et al.
Pubblicazione: (2025)
di: Huang, Tiansheng, et al.
Pubblicazione: (2025)
Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation
di: Schwartz, Daniel, et al.
Pubblicazione: (2025)
di: Schwartz, Daniel, et al.
Pubblicazione: (2025)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
di: Gu, Tianle, et al.
Pubblicazione: (2025)
di: Gu, Tianle, et al.
Pubblicazione: (2025)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
di: Hu, Xulin, et al.
Pubblicazione: (2026)
di: Hu, Xulin, et al.
Pubblicazione: (2026)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
di: Xing, Wenpeng, et al.
Pubblicazione: (2025)
PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
di: Rao, Akshaj Prashanth, et al.
Pubblicazione: (2025)
di: Rao, Akshaj Prashanth, et al.
Pubblicazione: (2025)
Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG
di: Singh, Inderjeet, et al.
Pubblicazione: (2026)
di: Singh, Inderjeet, et al.
Pubblicazione: (2026)
LLM in the Shell: Generative Honeypots
di: Sladić, Muris, et al.
Pubblicazione: (2023)
di: Sladić, Muris, et al.
Pubblicazione: (2023)
Federated In-Context LLM Agent Learning
di: Wu, Panlong, et al.
Pubblicazione: (2024)
di: Wu, Panlong, et al.
Pubblicazione: (2024)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
di: Zhu, Sicheng, et al.
Pubblicazione: (2024)
di: Zhu, Sicheng, et al.
Pubblicazione: (2024)
Policy-Invisible Violations in LLM-Based Agents
di: Wu, Jie, et al.
Pubblicazione: (2026)
di: Wu, Jie, et al.
Pubblicazione: (2026)
Certifying LLM Safety against Adversarial Prompting
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
di: Kumar, Aounon, et al.
Pubblicazione: (2023)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
di: Li, Jianwei, et al.
Pubblicazione: (2023)
di: Li, Jianwei, et al.
Pubblicazione: (2023)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
di: Halawi, Danny, et al.
Pubblicazione: (2024)
di: Halawi, Danny, et al.
Pubblicazione: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
di: Zymet, Jesse, et al.
Pubblicazione: (2026)
di: Zymet, Jesse, et al.
Pubblicazione: (2026)
LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning
di: Spracklen, Joseph, et al.
Pubblicazione: (2026)
di: Spracklen, Joseph, et al.
Pubblicazione: (2026)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
di: Liang, Buyun, et al.
Pubblicazione: (2025)
di: Liang, Buyun, et al.
Pubblicazione: (2025)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
di: Benjamin, Victoria, et al.
Pubblicazione: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
LLM Cyber Evaluations Don't Capture Real-World Risk
di: Lukošiūtė, Kamilė, et al.
Pubblicazione: (2025)
di: Lukošiūtė, Kamilė, et al.
Pubblicazione: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
di: Wang, Yifei, et al.
Pubblicazione: (2024)
di: Wang, Yifei, et al.
Pubblicazione: (2024)
Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
di: Wu, Xiaoyu, et al.
Pubblicazione: (2025)
di: Wu, Xiaoyu, et al.
Pubblicazione: (2025)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
di: Panfilov, Alexander, et al.
Pubblicazione: (2025)
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation
di: Roy, Joyjit, et al.
Pubblicazione: (2026)
di: Roy, Joyjit, et al.
Pubblicazione: (2026)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
di: Zhao, Andrew, et al.
Pubblicazione: (2025)
di: Zhao, Andrew, et al.
Pubblicazione: (2025)
SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting
di: Zhang, Hanxiu, et al.
Pubblicazione: (2025)
di: Zhang, Hanxiu, et al.
Pubblicazione: (2025)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
di: Dobre, David, et al.
Pubblicazione: (2025)
di: Dobre, David, et al.
Pubblicazione: (2025)
From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
di: Xu, Xiaoyu, et al.
Pubblicazione: (2026)
di: Xu, Xiaoyu, et al.
Pubblicazione: (2026)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
di: Cao, Bochuan, et al.
Pubblicazione: (2023)
A Fast, Reliable, and Secure Programming Language for LLM Agents with Code Actions
di: Mell, Stephen, et al.
Pubblicazione: (2025)
di: Mell, Stephen, et al.
Pubblicazione: (2025)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
di: Labiad, Ismail, et al.
Pubblicazione: (2025)
di: Labiad, Ismail, et al.
Pubblicazione: (2025)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
di: Zhang, Mohan, et al.
Pubblicazione: (2026)
Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates
di: Zheng, Xiaosen, et al.
Pubblicazione: (2024)
di: Zheng, Xiaosen, et al.
Pubblicazione: (2024)
There Are No Silly Questions: Evaluation of Offline LLM Capabilities from a Turkish Perspective
di: Yilmaz, Edibe, et al.
Pubblicazione: (2026)
di: Yilmaz, Edibe, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2026) -
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
di: Pawlak, Stanisław, et al.
Pubblicazione: (2025) -
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
di: Dubiński, Jan, et al.
Pubblicazione: (2026) -
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
di: Liang, Buyun, et al.
Pubblicazione: (2026) -
FLAME: Flexible LLM-Assisted Moderation Engine
di: Bakulin, Ivan, et al.
Pubblicazione: (2025)