Universal and Context-Independent Triggers for Precise Control of LLM Outputs
Fuente:
arXiv
Guardado en:
| Autores principales: | Liang, Jiashuo, Li, Guancheng, Yu, Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
por: Wang, Yanbo, et al.
Publicado: (2026)
por: Wang, Yanbo, et al.
Publicado: (2026)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
por: Xing, Wenpeng, et al.
Publicado: (2025)
por: Xing, Wenpeng, et al.
Publicado: (2025)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
por: Yu, Yongcan, et al.
Publicado: (2025)
por: Yu, Yongcan, et al.
Publicado: (2025)
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
por: Uenal, Fatih
Publicado: (2026)
por: Uenal, Fatih
Publicado: (2026)
CATMark: A Context-Aware Thresholding Framework for Robust Cross-Task Watermarking in Large Language Models
por: Zhang, Yu, et al.
Publicado: (2025)
por: Zhang, Yu, et al.
Publicado: (2025)
Federated In-Context LLM Agent Learning
por: Wu, Panlong, et al.
Publicado: (2024)
por: Wu, Panlong, et al.
Publicado: (2024)
An Independent Safety Evaluation of Kimi K2.5
por: Yong, Zheng-Xin, et al.
Publicado: (2026)
por: Yong, Zheng-Xin, et al.
Publicado: (2026)
Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems
por: Wang, Xiaoqing, et al.
Publicado: (2025)
por: Wang, Xiaoqing, et al.
Publicado: (2025)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
por: Wei, Jiali, et al.
Publicado: (2026)
por: Wei, Jiali, et al.
Publicado: (2026)
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
por: Liu, Junhao, et al.
Publicado: (2025)
por: Liu, Junhao, et al.
Publicado: (2025)
Searching for Privacy Risks in LLM Agents via Simulation
por: Zhang, Yanzhe, et al.
Publicado: (2025)
por: Zhang, Yanzhe, et al.
Publicado: (2025)
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
por: Kim, Jinhwa, et al.
Publicado: (2025)
por: Kim, Jinhwa, et al.
Publicado: (2025)
MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
por: Wang, Yuhui, et al.
Publicado: (2026)
por: Wang, Yuhui, et al.
Publicado: (2026)
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content
por: Guo, Ruoqi, et al.
Publicado: (2026)
por: Guo, Ruoqi, et al.
Publicado: (2026)
Contextualized Privacy Defense for LLM Agents
por: Wen, Yule, et al.
Publicado: (2026)
por: Wen, Yule, et al.
Publicado: (2026)
Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
por: Qu, Yubin, et al.
Publicado: (2026)
por: Qu, Yubin, et al.
Publicado: (2026)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
por: Yang, Tiankai, et al.
Publicado: (2026)
por: Yang, Tiankai, et al.
Publicado: (2026)
Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation
por: Shao, Shuo, et al.
Publicado: (2025)
por: Shao, Shuo, et al.
Publicado: (2025)
Token-level Data Selection for Safe LLM Fine-tuning
por: Li, Yanping, et al.
Publicado: (2026)
por: Li, Yanping, et al.
Publicado: (2026)
LLM Jailbreak Detection for (Almost) Free!
por: Chen, Guorui, et al.
Publicado: (2025)
por: Chen, Guorui, et al.
Publicado: (2025)
Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents
por: Liang, Zhibo, et al.
Publicado: (2025)
por: Liang, Zhibo, et al.
Publicado: (2025)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
por: Du, Wei, et al.
Publicado: (2023)
por: Du, Wei, et al.
Publicado: (2023)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
por: Yu, Miao, et al.
Publicado: (2024)
por: Yu, Miao, et al.
Publicado: (2024)
RedSage: A Cybersecurity Generalist LLM
por: Suryanto, Naufal, et al.
Publicado: (2026)
por: Suryanto, Naufal, et al.
Publicado: (2026)
Conversational Context Classification: A Representation Engineering Approach
por: Pan, Jonathan
Publicado: (2026)
por: Pan, Jonathan
Publicado: (2026)
LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation
por: Li, Xingyu, et al.
Publicado: (2025)
por: Li, Xingyu, et al.
Publicado: (2025)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
por: Yang, Wenkai, et al.
Publicado: (2024)
por: Yang, Wenkai, et al.
Publicado: (2024)
Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training
por: Yu, Yao-Ching, et al.
Publicado: (2025)
por: Yu, Yao-Ching, et al.
Publicado: (2025)
OneShield -- the Next Generation of LLM Guardrails
por: DeLuca, Chad, et al.
Publicado: (2025)
por: DeLuca, Chad, et al.
Publicado: (2025)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
por: Li, Xueyi, et al.
Publicado: (2026)
por: Li, Xueyi, et al.
Publicado: (2026)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
por: Li, Xirui, et al.
Publicado: (2024)
por: Li, Xirui, et al.
Publicado: (2024)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
por: Yoon, Sung-Hoon, et al.
Publicado: (2026)
por: Yoon, Sung-Hoon, et al.
Publicado: (2026)
LLM in the Shell: Generative Honeypots
por: Sladić, Muris, et al.
Publicado: (2023)
por: Sladić, Muris, et al.
Publicado: (2023)
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
por: Wang, Yanbo, et al.
Publicado: (2025)
por: Wang, Yanbo, et al.
Publicado: (2025)
CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
por: Zhou, Guanghao, et al.
Publicado: (2025)
por: Zhou, Guanghao, et al.
Publicado: (2025)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
por: Wang, Liwen, et al.
Publicado: (2025)
por: Wang, Liwen, et al.
Publicado: (2025)
Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings
por: Zhang, Yuanhe, et al.
Publicado: (2024)
por: Zhang, Yuanhe, et al.
Publicado: (2024)
AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection
por: Zhao, Zijie, et al.
Publicado: (2026)
por: Zhao, Zijie, et al.
Publicado: (2026)
FLAME: Flexible LLM-Assisted Moderation Engine
por: Bakulin, Ivan, et al.
Publicado: (2025)
por: Bakulin, Ivan, et al.
Publicado: (2025)
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
por: Xu, Huiyu, et al.
Publicado: (2024)
por: Xu, Huiyu, et al.
Publicado: (2024)
Ejemplares similares
-
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
por: Wang, Yanbo, et al.
Publicado: (2026) -
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
por: Xing, Wenpeng, et al.
Publicado: (2025) -
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
por: Yu, Yongcan, et al.
Publicado: (2025) -
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
por: Uenal, Fatih
Publicado: (2026) -
CATMark: A Context-Aware Thresholding Framework for Robust Cross-Task Watermarking in Large Language Models
por: Zhang, Yu, et al.
Publicado: (2025)