sudo rm -rf agentic_security
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Sejin, Kim, Jian, Park, Haon, Yousefpour, Ashkan, Yu, Sangyoon, Song, Min |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
by: Panpatil, Siddhant, et al.
Published: (2025)
by: Panpatil, Siddhant, et al.
Published: (2025)
M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs
by: Ha, Junwoo, et al.
Published: (2025)
by: Ha, Junwoo, et al.
Published: (2025)
sudoLLM: On Multi-role Alignment of Language Models
by: Saha, Soumadeep, et al.
Published: (2025)
by: Saha, Soumadeep, et al.
Published: (2025)
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026)
by: Lee, Hyomin, et al.
Published: (2026)
X-Teaming Evolutionary M2S: Automated Discovery of Multi-turn to Single-turn Jailbreak Templates
by: Kim, Hyunjun, et al.
Published: (2025)
by: Kim, Hyunjun, et al.
Published: (2025)
Claim-Guided Textual Backdoor Attack for Practical Applications
by: Song, Minkyoo, et al.
Published: (2024)
by: Song, Minkyoo, et al.
Published: (2024)
SGuard-v1: Safety Guardrail for Large Language Models
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
MURMUR: Using cross-user chatter to break collaborative language agents in groups
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
by: Wu, Yiran, et al.
Published: (2025)
by: Wu, Yiran, et al.
Published: (2025)
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization
by: Liu, Mingshuo, et al.
Published: (2026)
by: Liu, Mingshuo, et al.
Published: (2026)
SecureBreak -- A dataset towards safe and secure models
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
by: Yan, Yu, et al.
Published: (2025)
by: Yan, Yu, et al.
Published: (2025)
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
by: Yan, Dong, et al.
Published: (2026)
by: Yan, Dong, et al.
Published: (2026)
Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
by: Min, Nay Myat, et al.
Published: (2026)
by: Min, Nay Myat, et al.
Published: (2026)
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
by: Leong, Chak Tou, et al.
Published: (2025)
by: Leong, Chak Tou, et al.
Published: (2025)
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
by: Kim, Jinhwa, et al.
Published: (2025)
by: Kim, Jinhwa, et al.
Published: (2025)
Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position
by: Xie, Zhixin, et al.
Published: (2025)
by: Xie, Zhixin, et al.
Published: (2025)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
by: Chae, Kyubyung, et al.
Published: (2025)
by: Chae, Kyubyung, et al.
Published: (2025)
RedacBench: Can AI Erase Your Secrets?
by: Jeon, Hyunjun, et al.
Published: (2026)
by: Jeon, Hyunjun, et al.
Published: (2026)
Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
by: Gong, Bofan, et al.
Published: (2025)
by: Gong, Bofan, et al.
Published: (2025)
DMFI: A Dual-Modality Log Analysis Framework for Insider Threat Detection with LoRA-Tuned Language Models
by: Kong, Kaichuan, et al.
Published: (2025)
by: Kong, Kaichuan, et al.
Published: (2025)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
by: Jones, Jaylen, et al.
Published: (2026)
by: Jones, Jaylen, et al.
Published: (2026)
Direct Token Optimization: A Self-contained Approach to Large Language Model Unlearning
by: Lee, Hong kyu, et al.
Published: (2025)
by: Lee, Hong kyu, et al.
Published: (2025)
PRISON: Unmasking the Criminal Potential of Large Language Models
by: Wu, Xinyi, et al.
Published: (2025)
by: Wu, Xinyi, et al.
Published: (2025)
Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding
by: Qiu, Huming, et al.
Published: (2024)
by: Qiu, Huming, et al.
Published: (2024)
Beyond the Tip of Efficiency: Uncovering the Submerged Threats of Jailbreak Attacks in Small Language Models
by: Yi, Sibo, et al.
Published: (2025)
by: Yi, Sibo, et al.
Published: (2025)
Dagger Behind Smile: Fool LLMs with a Happy Ending Story
by: Song, Xurui, et al.
Published: (2025)
by: Song, Xurui, et al.
Published: (2025)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
by: Wang, Xilong, et al.
Published: (2026)
by: Wang, Xilong, et al.
Published: (2026)
Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models
by: Yu, Ye, et al.
Published: (2026)
by: Yu, Ye, et al.
Published: (2026)
Representation Bending for Large Language Model Safety
by: Yousefpour, Ashkan, et al.
Published: (2025)
by: Yousefpour, Ashkan, et al.
Published: (2025)
Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
by: Na, CheolWon, et al.
Published: (2025)
by: Na, CheolWon, et al.
Published: (2025)
Say Something Else: Rethinking Contextual Privacy as Information Sufficiency
by: Xiao, Yunze, et al.
Published: (2026)
by: Xiao, Yunze, et al.
Published: (2026)
Adversarial Attacks Against Automated Fact-Checking: A Survey
by: Liu, Fanzhen, et al.
Published: (2025)
by: Liu, Fanzhen, et al.
Published: (2025)
Building Intelligence Identification System via Large Language Model Watermarking: A Survey and Beyond
by: Wang, Xuhong, et al.
Published: (2024)
by: Wang, Xuhong, et al.
Published: (2024)
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
by: Liang, Jiashuo, et al.
Published: (2024)
by: Liang, Jiashuo, et al.
Published: (2024)
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
by: Zhou, Kaiwen, et al.
Published: (2025)
by: Zhou, Kaiwen, et al.
Published: (2025)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Similar Items
-
Eliciting and Analyzing Emergent Misalignment in State-of-the-Art Large Language Models
by: Panpatil, Siddhant, et al.
Published: (2025) -
M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs
by: Ha, Junwoo, et al.
Published: (2025) -
sudoLLM: On Multi-role Alignment of Language Models
by: Saha, Soumadeep, et al.
Published: (2025) -
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
by: Lee, Hyomin, et al.
Published: (2026) -
X-Teaming Evolutionary M2S: Automated Discovery of Multi-turn to Single-turn Jailbreak Templates
by: Kim, Hyunjun, et al.
Published: (2025)