FoolSDEdit: Deceptively Steering Your Edits Towards Targeted Attribute-aware Distribution
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Qi, Wang, Dongxia, Li, Tianlin, Xu, Zhihong, Liu, Yang, Ren, Kui, Wang, Wenhai, Guo, Qing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dynamic Deception: When Pedestrians Team Up to Fool Autonomous Cars
por: Tehrani, Masoud Jamshidiyan, et al.
Publicado: (2026)
por: Tehrani, Masoud Jamshidiyan, et al.
Publicado: (2026)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
por: Yuan, Xiaohan, et al.
Publicado: (2024)
por: Yuan, Xiaohan, et al.
Publicado: (2024)
Defending LVLMs Against Vision Attacks through Partial-Perception Supervision
por: Zhou, Qi, et al.
Publicado: (2024)
por: Zhou, Qi, et al.
Publicado: (2024)
Resource-aware Cyber Deception for Microservice-based Applications
por: Zambianco, Marco, et al.
Publicado: (2023)
por: Zambianco, Marco, et al.
Publicado: (2023)
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy
por: Jeong, Joonhyun, et al.
Publicado: (2025)
por: Jeong, Joonhyun, et al.
Publicado: (2025)
On the Robustness of LDP Protocols for Numerical Attributes under Data Poisoning Attacks
por: Li, Xiaoguang, et al.
Publicado: (2024)
por: Li, Xiaoguang, et al.
Publicado: (2024)
BadEdit: Backdooring large language models by model editing
por: Li, Yanzhou, et al.
Publicado: (2024)
por: Li, Yanzhou, et al.
Publicado: (2024)
PT-Mark: Invisible Watermarking for Text-to-image Diffusion Models via Semantic-aware Pivotal Tuning
por: Wang, Yaopeng, et al.
Publicado: (2025)
por: Wang, Yaopeng, et al.
Publicado: (2025)
Mirage Fools the Ear, Mute Hides the Truth: Precise Targeted Adversarial Attacks on Polyphonic Sound Event Detection Systems
por: Su, Junjie, et al.
Publicado: (2025)
por: Su, Junjie, et al.
Publicado: (2025)
EditMF: Drawing an Invisible Fingerprint for Your Large Language Models
por: Wu, Jiaxuan, et al.
Publicado: (2025)
por: Wu, Jiaxuan, et al.
Publicado: (2025)
StyleFool: Fooling Video Classification Systems via Style Transfer
por: Cao, Yuxin, et al.
Publicado: (2022)
por: Cao, Yuxin, et al.
Publicado: (2022)
Evolving Deception: When Agents Evolve, Deception Wins
por: Ying, Zonghao, et al.
Publicado: (2026)
por: Ying, Zonghao, et al.
Publicado: (2026)
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
por: Qi, Weiwei, et al.
Publicado: (2026)
por: Qi, Weiwei, et al.
Publicado: (2026)
WiP: Deception-in-Depth Using Multiple Layers of Deception
por: Landsborough, Jason, et al.
Publicado: (2024)
por: Landsborough, Jason, et al.
Publicado: (2024)
Towards Personal Data Sharing Autonomy:A Task-driven Data Capsule Sharing System
por: Lyu, Qiuyun, et al.
Publicado: (2024)
por: Lyu, Qiuyun, et al.
Publicado: (2024)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
por: Zhao, Weixiang, et al.
Publicado: (2025)
por: Zhao, Weixiang, et al.
Publicado: (2025)
Target Attack Backdoor Malware Analysis and Attribution
por: Lai, Anthony Cheuk Tung, et al.
Publicado: (2025)
por: Lai, Anthony Cheuk Tung, et al.
Publicado: (2025)
Towards in-situ Psychological Profiling of Cybercriminals Using Dynamically Generated Deception Environments
por: Quibell, Jacob
Publicado: (2024)
por: Quibell, Jacob
Publicado: (2024)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
por: Fu, Yuchuan, et al.
Publicado: (2025)
por: Fu, Yuchuan, et al.
Publicado: (2025)
Your Trust, Your Terms: A General Paradigm for Near-Instant Cross-Chain Transfer
por: Wu, Di, et al.
Publicado: (2024)
por: Wu, Di, et al.
Publicado: (2024)
Mitigating Data Poisoning Attacks to Local Differential Privacy
por: Li, Xiaolin, et al.
Publicado: (2025)
por: Li, Xiaolin, et al.
Publicado: (2025)
I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps
por: Cai, Yifeng, et al.
Publicado: (2025)
por: Cai, Yifeng, et al.
Publicado: (2025)
Can LLMs be Fooled? Investigating Vulnerabilities in LLMs
por: Abdali, Sara, et al.
Publicado: (2024)
por: Abdali, Sara, et al.
Publicado: (2024)
A Cascade Approach for APT Campaign Attribution in System Event Logs: Technique Hunting and Subgraph Matching
por: Huang, Yi-Ting, et al.
Publicado: (2024)
por: Huang, Yi-Ting, et al.
Publicado: (2024)
Protecting Your Voice: Temporal-aware Robust Watermarking
por: Li, Yue, et al.
Publicado: (2025)
por: Li, Yue, et al.
Publicado: (2025)
Dynamic Target Attack
por: Xiu, Kedong, et al.
Publicado: (2025)
por: Xiu, Kedong, et al.
Publicado: (2025)
Toward Unbiased Multiple-Target Fuzzing with Path Diversity
por: Rong, Huanyao, et al.
Publicado: (2023)
por: Rong, Huanyao, et al.
Publicado: (2023)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
por: Shao, Shuo, et al.
Publicado: (2024)
por: Shao, Shuo, et al.
Publicado: (2024)
Towards Better Attribute Inference Vulnerability Measures
por: Francis, Paul, et al.
Publicado: (2025)
por: Francis, Paul, et al.
Publicado: (2025)
Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
por: Guo, Xuyang, et al.
Publicado: (2025)
por: Guo, Xuyang, et al.
Publicado: (2025)
High-Rate Public-Key Pseudorandom Codes for Edit Errors
por: Huang, Shengtang, et al.
Publicado: (2026)
por: Huang, Shengtang, et al.
Publicado: (2026)
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
por: Zeng, Xiyu, et al.
Publicado: (2025)
por: Zeng, Xiyu, et al.
Publicado: (2025)
FINER: Enhancing State-of-the-art Classifiers with Feature Attribution to Facilitate Security Analysis
por: He, Yiling, et al.
Publicado: (2023)
por: He, Yiling, et al.
Publicado: (2023)
FlowSteer: Prompt-Only Workflow Steering Exposes Planning-Time Vulnerabilities in Multi-Agent LLM Systems
por: Li, Fanxiao, et al.
Publicado: (2026)
por: Li, Fanxiao, et al.
Publicado: (2026)
I Know Who Clones Your Code: Interpretable Smart Contract Similarity Detection
por: Liu, Zhenguang, et al.
Publicado: (2025)
por: Liu, Zhenguang, et al.
Publicado: (2025)
Enhancing Distributed Authorization With Lagrange Interpolation And Attribute-Based Encryption
por: Sinha, Keshav, et al.
Publicado: (2025)
por: Sinha, Keshav, et al.
Publicado: (2025)
When and How to Fool Explainable Models (and Humans) with Adversarial Examples
por: Vadillo, Jon, et al.
Publicado: (2021)
por: Vadillo, Jon, et al.
Publicado: (2021)
EXTree: Towards Supporting Explainability in Attribute-based Access Control
por: Chowdary, Shanampudi Pranaya, et al.
Publicado: (2026)
por: Chowdary, Shanampudi Pranaya, et al.
Publicado: (2026)
Illusion Worlds: Deceptive UI Attacks in Social VR
por: Lee, Junhee, et al.
Publicado: (2025)
por: Lee, Junhee, et al.
Publicado: (2025)
Koney: A Cyber Deception Orchestration Framework for Kubernetes
por: Kahlhofer, Mario, et al.
Publicado: (2025)
por: Kahlhofer, Mario, et al.
Publicado: (2025)
Ejemplares similares
-
Dynamic Deception: When Pedestrians Team Up to Fool Autonomous Cars
por: Tehrani, Masoud Jamshidiyan, et al.
Publicado: (2026) -
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
por: Yuan, Xiaohan, et al.
Publicado: (2024) -
Defending LVLMs Against Vision Attacks through Partial-Perception Supervision
por: Zhou, Qi, et al.
Publicado: (2024) -
Resource-aware Cyber Deception for Microservice-based Applications
por: Zambianco, Marco, et al.
Publicado: (2023) -
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy
por: Jeong, Joonhyun, et al.
Publicado: (2025)