EnsembleSHAP: Faithful and Certifiably Robust Attribution for Random Subspace Method
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Yanting, Jia, Jinyuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FCert: Certifiably Robust Few-Shot Classification in the Era of Foundation Models
por: Wang, Yanting, et al.
Publicado: (2024)
por: Wang, Yanting, et al.
Publicado: (2024)
AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
por: Wang, Yanting, et al.
Publicado: (2025)
por: Wang, Yanting, et al.
Publicado: (2025)
TracLLM: A Generic Framework for Attributing Long Context LLMs
por: Wang, Yanting, et al.
Publicado: (2025)
por: Wang, Yanting, et al.
Publicado: (2025)
AgentWatcher: A Rule-based Prompt Injection Monitor
por: Wang, Yanting, et al.
Publicado: (2026)
por: Wang, Yanting, et al.
Publicado: (2026)
FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
por: Wang, Yanting, et al.
Publicado: (2026)
por: Wang, Yanting, et al.
Publicado: (2026)
TrojanDec: Data-free Detection of Trojan Inputs in Self-supervised Learning
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
Distributed Backdoor Attacks on Federated Graph Learning and Certified Defenses
por: Yang, Yuxin, et al.
Publicado: (2024)
por: Yang, Yuxin, et al.
Publicado: (2024)
TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
por: Wang, Yanting, et al.
Publicado: (2025)
por: Wang, Yanting, et al.
Publicado: (2025)
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
por: Yin, Chenlong, et al.
Publicado: (2026)
por: Yin, Chenlong, et al.
Publicado: (2026)
Certifiably Robust Image Watermark
por: Jiang, Zhengyuan, et al.
Publicado: (2024)
por: Jiang, Zhengyuan, et al.
Publicado: (2024)
SecInfer: Preventing Prompt Injection via Inference-time Scaling
por: Liu, Yupei, et al.
Publicado: (2025)
por: Liu, Yupei, et al.
Publicado: (2025)
MMCert: Provable Defense against Adversarial Attacks to Multi-modal Models
por: Wang, Yanting, et al.
Publicado: (2024)
por: Wang, Yanting, et al.
Publicado: (2024)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
por: Geng, Runpeng, et al.
Publicado: (2025)
por: Geng, Runpeng, et al.
Publicado: (2025)
Provably Robust Explainable Graph Neural Networks against Graph Perturbation Attacks
por: Li, Jiate, et al.
Publicado: (2025)
por: Li, Jiate, et al.
Publicado: (2025)
PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
por: Zou, Wei, et al.
Publicado: (2025)
por: Zou, Wei, et al.
Publicado: (2025)
RobustMask: Certified Robustness against Adversarial Neural Ranking Attack via Randomized Masking
por: Liu, Jiawei, et al.
Publicado: (2025)
por: Liu, Jiawei, et al.
Publicado: (2025)
Ideal Attribution and Faithful Watermarks for Language Models
por: Song, Min Jae, et al.
Publicado: (2025)
por: Song, Min Jae, et al.
Publicado: (2025)
Adaptive Randomized Smoothing: Certified Adversarial Robustness for Multi-Step Defences
por: Lyu, Saiyue, et al.
Publicado: (2024)
por: Lyu, Saiyue, et al.
Publicado: (2024)
FedGMark: Certifiably Robust Watermarking for Federated Graph Learning
por: Yang, Yuxin, et al.
Publicado: (2024)
por: Yang, Yuxin, et al.
Publicado: (2024)
Registered Attribute-Based Encryption with Publicly Verifiable Certified Deletion, Everlasting Security, and More
por: Murshid, Shayeef, et al.
Publicado: (2026)
por: Murshid, Shayeef, et al.
Publicado: (2026)
Interpretable Ensemble Learning for Network Traffic Anomaly Detection: A SHAP-based Explainable AI Framework for Embedded Systems Security
por: Shao, Wanru
Publicado: (2026)
por: Shao, Wanru
Publicado: (2026)
Explainable Threat Attribution for IoT Networks Using Conditional SHAP and Flow Behavior Modelling
por: Ozechi, Samuel, et al.
Publicado: (2026)
por: Ozechi, Samuel, et al.
Publicado: (2026)
A Certified Robust Watermark For Large Language Models
por: Feng, Xianheng, et al.
Publicado: (2024)
por: Feng, Xianheng, et al.
Publicado: (2024)
A Robust Certified Machine Unlearning Method Under Distribution Shift
por: Guo, Jinduo, et al.
Publicado: (2026)
por: Guo, Jinduo, et al.
Publicado: (2026)
PIArena: A Platform for Prompt Injection Evaluation
por: Geng, Runpeng, et al.
Publicado: (2026)
por: Geng, Runpeng, et al.
Publicado: (2026)
Towards Strong Certified Defense with Universal Asymmetric Randomization
por: Hong, Hanbin, et al.
Publicado: (2025)
por: Hong, Hanbin, et al.
Publicado: (2025)
Boosting Certified Robustness for Time Series Classification with Efficient Self-Ensemble
por: Dong, Chang, et al.
Publicado: (2024)
por: Dong, Chang, et al.
Publicado: (2024)
Evaluating LLM-based Personal Information Extraction and Countermeasures
por: Liu, Yupei, et al.
Publicado: (2024)
por: Liu, Yupei, et al.
Publicado: (2024)
Certified PEFTSmoothing: Parameter-Efficient Fine-Tuning with Randomized Smoothing
por: Fu, Chengyan, et al.
Publicado: (2024)
por: Fu, Chengyan, et al.
Publicado: (2024)
Certified Adversarial Robustness of Machine Learning-based Malware Detectors via (De)Randomized Smoothing
por: Gibert, Daniel, et al.
Publicado: (2024)
por: Gibert, Daniel, et al.
Publicado: (2024)
Getting a-Round Guarantees: Floating-Point Attacks on Certified Robustness
por: Jin, Jiankai, et al.
Publicado: (2022)
por: Jin, Jiankai, et al.
Publicado: (2022)
AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
por: Huang, Zhuoqun, et al.
Publicado: (2025)
por: Huang, Zhuoqun, et al.
Publicado: (2025)
TRUSTDESC: Preventing Tool Poisoning in LLM Applications via Trusted Description Generation
por: Ye, Hengkai, et al.
Publicado: (2026)
por: Ye, Hengkai, et al.
Publicado: (2026)
Provably Robust Multi-bit Watermarking for AI-generated Text
por: Qu, Wenjie, et al.
Publicado: (2024)
por: Qu, Wenjie, et al.
Publicado: (2024)
On the Equivalence between Classical Position Verification and Certified Randomness
por: Kaleoglu, Fatih, et al.
Publicado: (2024)
por: Kaleoglu, Fatih, et al.
Publicado: (2024)
Privacy-Optimized Randomized Response for Sharing Multi-Attribute Data
por: Yamamoto, Akito, et al.
Publicado: (2024)
por: Yamamoto, Akito, et al.
Publicado: (2024)
On Using Certified Training towards Empirical Robustness
por: De Palma, Alessandro, et al.
Publicado: (2024)
por: De Palma, Alessandro, et al.
Publicado: (2024)
Certified Causal Attribution for Real-Time Attack Forensics in 6G Network Slicing
por: Quan, Minh K., et al.
Publicado: (2026)
por: Quan, Minh K., et al.
Publicado: (2026)
PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization
por: Geng, Runpeng, et al.
Publicado: (2025)
por: Geng, Runpeng, et al.
Publicado: (2025)
Interpretable Anomaly Detection in Encrypted Traffic Using SHAP with Machine Learning Models
por: Singh, Kalindi, et al.
Publicado: (2025)
por: Singh, Kalindi, et al.
Publicado: (2025)
Ejemplares similares
-
FCert: Certifiably Robust Few-Shot Classification in the Era of Foundation Models
por: Wang, Yanting, et al.
Publicado: (2024) -
AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
por: Wang, Yanting, et al.
Publicado: (2025) -
TracLLM: A Generic Framework for Attributing Long Context LLMs
por: Wang, Yanting, et al.
Publicado: (2025) -
AgentWatcher: A Rule-based Prompt Injection Monitor
por: Wang, Yanting, et al.
Publicado: (2026) -
FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption
por: Wang, Yanting, et al.
Publicado: (2026)