SALLIE: Safeguarding Against Latent Language & Image Exploits
Fuente:
arXiv
Salvato in:
| Autori principali: | Azov, Guy, Rivlin, Ofer, Shtar, Guy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
di: Du, Guanchen, et al.
Pubblicazione: (2025)
di: Du, Guanchen, et al.
Pubblicazione: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
di: Dang, Kieu, et al.
Pubblicazione: (2025)
di: Dang, Kieu, et al.
Pubblicazione: (2025)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
di: Young, Richard J., et al.
Pubblicazione: (2026)
di: Young, Richard J., et al.
Pubblicazione: (2026)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
di: Morales, Jaime, et al.
Pubblicazione: (2026)
di: Morales, Jaime, et al.
Pubblicazione: (2026)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
di: Gu, Yongtong, et al.
Pubblicazione: (2026)
di: Gu, Yongtong, et al.
Pubblicazione: (2026)
Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats
di: Zhang, Bingxue, et al.
Pubblicazione: (2026)
di: Zhang, Bingxue, et al.
Pubblicazione: (2026)
Sensitivity Uncertainty Alignment in Large Language Models
di: Hiremath, Prakul Sunil, et al.
Pubblicazione: (2026)
di: Hiremath, Prakul Sunil, et al.
Pubblicazione: (2026)
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
di: Doda, Shravan
Pubblicazione: (2026)
di: Doda, Shravan
Pubblicazione: (2026)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
di: Othman, Refat
Pubblicazione: (2026)
di: Othman, Refat
Pubblicazione: (2026)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
di: Lelle, Travis
Pubblicazione: (2026)
di: Lelle, Travis
Pubblicazione: (2026)
Semantically Guided Adversarial Testing of Vision Models Using Language Models
di: Filus, Katarzyna, et al.
Pubblicazione: (2025)
di: Filus, Katarzyna, et al.
Pubblicazione: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
di: Portelance, Eva, et al.
Pubblicazione: (2023)
di: Portelance, Eva, et al.
Pubblicazione: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
di: Yang, Shan
Pubblicazione: (2026)
di: Yang, Shan
Pubblicazione: (2026)
Retrieval Augmented Classification for Confidential Documents
di: Chang, Yeseul E., et al.
Pubblicazione: (2026)
di: Chang, Yeseul E., et al.
Pubblicazione: (2026)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
di: Dawson, Ads, et al.
Pubblicazione: (2025)
di: Dawson, Ads, et al.
Pubblicazione: (2025)
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
di: Wu, Shuai, et al.
Pubblicazione: (2026)
di: Wu, Shuai, et al.
Pubblicazione: (2026)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
di: Yu, Hao, et al.
Pubblicazione: (2026)
di: Yu, Hao, et al.
Pubblicazione: (2026)
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
di: Hoang, Tien Dat
Pubblicazione: (2025)
di: Hoang, Tien Dat
Pubblicazione: (2025)
Countermind: A Multi-Layered Security Architecture for Large Language Models
di: Schwarz, Dominik
Pubblicazione: (2025)
di: Schwarz, Dominik
Pubblicazione: (2025)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
di: Zanbaghi, Shahin, et al.
Pubblicazione: (2025)
di: Zanbaghi, Shahin, et al.
Pubblicazione: (2025)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
di: Bercovich, Ivan, et al.
Pubblicazione: (2026)
di: Bercovich, Ivan, et al.
Pubblicazione: (2026)
The Automation Advantage in AI Red Teaming
di: Mulla, Rob, et al.
Pubblicazione: (2025)
di: Mulla, Rob, et al.
Pubblicazione: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
di: Amini, Ali
Pubblicazione: (2025)
di: Amini, Ali
Pubblicazione: (2025)
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
di: Usman, Rana Muhammad
Pubblicazione: (2026)
di: Usman, Rana Muhammad
Pubblicazione: (2026)
Semantic Superiority vs. Forensic Efficiency: A Comparative Analysis of Deep Learning and Psycholinguistics for Business Email Compromise Detection
di: Adjei, Yaw Osei, et al.
Pubblicazione: (2025)
di: Adjei, Yaw Osei, et al.
Pubblicazione: (2025)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
di: Nathanson, Samuel, et al.
Pubblicazione: (2025)
di: Nathanson, Samuel, et al.
Pubblicazione: (2025)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
di: Hill, Brennen, et al.
Pubblicazione: (2025)
di: Hill, Brennen, et al.
Pubblicazione: (2025)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
di: Wang, Haochuan Kevin, et al.
Pubblicazione: (2026)
di: Wang, Haochuan Kevin, et al.
Pubblicazione: (2026)
Towards Modeling Cybersecurity Behavior of Humans in Organizations
di: Kürtz, Klaas Ole
Pubblicazione: (2026)
di: Kürtz, Klaas Ole
Pubblicazione: (2026)
Send to which account? Evaluation of an LLM-based Scambaiting System
di: Siadati, Hossein, et al.
Pubblicazione: (2025)
di: Siadati, Hossein, et al.
Pubblicazione: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
di: Chona, Alankrit, et al.
Pubblicazione: (2026)
di: Chona, Alankrit, et al.
Pubblicazione: (2026)
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts
di: Young, Richard J., et al.
Pubblicazione: (2026)
di: Young, Richard J., et al.
Pubblicazione: (2026)
Measuring Harmfulness of Computer-Using Agents
di: Tian, Aaron Xuxiang, et al.
Pubblicazione: (2025)
di: Tian, Aaron Xuxiang, et al.
Pubblicazione: (2025)
Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack
di: Abdukhamidov, Eldor, et al.
Pubblicazione: (2025)
di: Abdukhamidov, Eldor, et al.
Pubblicazione: (2025)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
di: Ge, Yuxu
Pubblicazione: (2026)
di: Ge, Yuxu
Pubblicazione: (2026)
Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning
di: Subedar, Noah, et al.
Pubblicazione: (2025)
di: Subedar, Noah, et al.
Pubblicazione: (2025)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
di: Chakraborty, Amit, et al.
Pubblicazione: (2025)
di: Chakraborty, Amit, et al.
Pubblicazione: (2025)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
di: Seddik, Issam, et al.
Pubblicazione: (2025)
di: Seddik, Issam, et al.
Pubblicazione: (2025)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
di: Portelance, Eva, et al.
Pubblicazione: (2024)
di: Portelance, Eva, et al.
Pubblicazione: (2024)
Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility
di: Engelberg, Gal, et al.
Pubblicazione: (2026)
di: Engelberg, Gal, et al.
Pubblicazione: (2026)
Documenti analoghi
-
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
di: Du, Guanchen, et al.
Pubblicazione: (2025) -
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
di: Dang, Kieu, et al.
Pubblicazione: (2025) -
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
di: Young, Richard J., et al.
Pubblicazione: (2026) -
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
di: Morales, Jaime, et al.
Pubblicazione: (2026) -
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
di: Gu, Yongtong, et al.
Pubblicazione: (2026)