Subversion via Focal Points: Investigating Collusion in LLM Monitoring
Fuente:
arXiv
Salvato in:
| Autore principale: | Järviniemi, Olli |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
di: Mathew, Yohan, et al.
Pubblicazione: (2024)
di: Mathew, Yohan, et al.
Pubblicazione: (2024)
Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference
di: Xue, Anton, et al.
Pubblicazione: (2024)
di: Xue, Anton, et al.
Pubblicazione: (2024)
Interpretable LLM Guardrails via Sparse Representation Steering
di: He, Zeqing, et al.
Pubblicazione: (2025)
di: He, Zeqing, et al.
Pubblicazione: (2025)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
di: Horal, Artur, et al.
Pubblicazione: (2025)
di: Horal, Artur, et al.
Pubblicazione: (2025)
MaskSQL: Safeguarding Privacy for LLM-Based Text-to-SQL via Abstraction
di: Abedini, Sepideh, et al.
Pubblicazione: (2025)
di: Abedini, Sepideh, et al.
Pubblicazione: (2025)
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
di: Jawad, Huseein, et al.
Pubblicazione: (2025)
di: Jawad, Huseein, et al.
Pubblicazione: (2025)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
di: Zhang, Xiaozhe, et al.
Pubblicazione: (2026)
di: Zhang, Xiaozhe, et al.
Pubblicazione: (2026)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
di: Li, Yucheng, et al.
Pubblicazione: (2025)
di: Li, Yucheng, et al.
Pubblicazione: (2025)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
di: Hui, Zheng, et al.
Pubblicazione: (2025)
di: Hui, Zheng, et al.
Pubblicazione: (2025)
Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI
di: Sandoval, Aaron, et al.
Pubblicazione: (2025)
di: Sandoval, Aaron, et al.
Pubblicazione: (2025)
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2026)
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2026)
PRISM: Privacy-Aware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch Collaboration
di: Zhan, Junfei, et al.
Pubblicazione: (2025)
di: Zhan, Junfei, et al.
Pubblicazione: (2025)
Enabling Efficient Attack Investigation via Human-in-the-Loop Security Analysis
di: Tsegai, Saimon Amanuel, et al.
Pubblicazione: (2022)
di: Tsegai, Saimon Amanuel, et al.
Pubblicazione: (2022)
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
di: German, Eyal, et al.
Pubblicazione: (2025)
di: German, Eyal, et al.
Pubblicazione: (2025)
LLM Reinforcement in Context
di: Rivasseau, Thomas
Pubblicazione: (2025)
di: Rivasseau, Thomas
Pubblicazione: (2025)
Watermarking LLM Agent Trajectories
di: Meng, Wenlong, et al.
Pubblicazione: (2026)
di: Meng, Wenlong, et al.
Pubblicazione: (2026)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
di: Wu, Yiran, et al.
Pubblicazione: (2025)
di: Wu, Yiran, et al.
Pubblicazione: (2025)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
di: Yang, Wenkai, et al.
Pubblicazione: (2024)
di: Yang, Wenkai, et al.
Pubblicazione: (2024)
Proactive defense against LLM Jailbreak
di: Zhao, Weiliang, et al.
Pubblicazione: (2025)
di: Zhao, Weiliang, et al.
Pubblicazione: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
di: Chen, Bocheng, et al.
Pubblicazione: (2024)
di: Chen, Bocheng, et al.
Pubblicazione: (2024)
LLM Anonymization Against Agentic Re-Identification
di: Li, Ziwen, et al.
Pubblicazione: (2026)
di: Li, Ziwen, et al.
Pubblicazione: (2026)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
di: Freenor, Michael, et al.
Pubblicazione: (2025)
di: Freenor, Michael, et al.
Pubblicazione: (2025)
Multi-use LLM Watermarking and the False Detection Problem
di: Fu, Zihao, et al.
Pubblicazione: (2025)
di: Fu, Zihao, et al.
Pubblicazione: (2025)
FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework
di: Béjar, Mario Rodríguez, et al.
Pubblicazione: (2026)
di: Béjar, Mario Rodríguez, et al.
Pubblicazione: (2026)
On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning
di: Ye, Xiaotian, et al.
Pubblicazione: (2026)
di: Ye, Xiaotian, et al.
Pubblicazione: (2026)
Security Attacks on LLM-based Code Completion Tools
di: Cheng, Wen, et al.
Pubblicazione: (2024)
di: Cheng, Wen, et al.
Pubblicazione: (2024)
GLiGuard: Schema-Conditioned Classification for LLM Safeguard
di: Zaratiana, Urchade, et al.
Pubblicazione: (2026)
di: Zaratiana, Urchade, et al.
Pubblicazione: (2026)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
di: Li, Caihua, et al.
Pubblicazione: (2024)
di: Li, Caihua, et al.
Pubblicazione: (2024)
WorldCup Sampling for Multi-bit LLM Watermarking
di: Wang, Yidan, et al.
Pubblicazione: (2026)
di: Wang, Yidan, et al.
Pubblicazione: (2026)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
di: Wang, Junlin, et al.
Pubblicazione: (2024)
di: Wang, Junlin, et al.
Pubblicazione: (2024)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
di: Wu, Yu-Hang, et al.
Pubblicazione: (2025)
di: Wu, Yu-Hang, et al.
Pubblicazione: (2025)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
di: Fu, Wenjie, et al.
Pubblicazione: (2026)
di: Fu, Wenjie, et al.
Pubblicazione: (2026)
Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
di: Zeng, Xinyi, et al.
Pubblicazione: (2024)
di: Zeng, Xinyi, et al.
Pubblicazione: (2024)
Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection
di: Dang, Kieu, et al.
Pubblicazione: (2026)
di: Dang, Kieu, et al.
Pubblicazione: (2026)
"Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers
di: Zhou, Qin, et al.
Pubblicazione: (2025)
di: Zhou, Qin, et al.
Pubblicazione: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
di: Shen, Guobin, et al.
Pubblicazione: (2025)
di: Shen, Guobin, et al.
Pubblicazione: (2025)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
di: Hastuti, Rochana Prih, et al.
Pubblicazione: (2025)
di: Hastuti, Rochana Prih, et al.
Pubblicazione: (2025)
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
di: Liu, Yepeng, et al.
Pubblicazione: (2025)
di: Liu, Yepeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
di: Mathew, Yohan, et al.
Pubblicazione: (2024) -
Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference
di: Xue, Anton, et al.
Pubblicazione: (2024) -
Interpretable LLM Guardrails via Sparse Representation Steering
di: He, Zeqing, et al.
Pubblicazione: (2025) -
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
di: Horal, Artur, et al.
Pubblicazione: (2025) -
MaskSQL: Safeguarding Privacy for LLM-Based Text-to-SQL via Abstraction
di: Abedini, Sepideh, et al.
Pubblicazione: (2025)