Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information
Fuente:
arXiv
Guardado en:
| Autores principales: | Joung, Youngju, Lee, Sehyun, Choi, Jaesik |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BugSweeper: Function-Level Detection of Smart Contract Vulnerabilities Using Graph Neural Networks
por: Lee, Uisang, et al.
Publicado: (2025)
por: Lee, Uisang, et al.
Publicado: (2025)
Unveiling the Threat of Fraud Gangs to Graph Neural Networks: Multi-Target Graph Injection Attacks Against GNN-Based Fraud Detectors
por: Choi, Jinhyeok, et al.
Publicado: (2024)
por: Choi, Jinhyeok, et al.
Publicado: (2024)
Statement-Level Vulnerability Detection: Learning Vulnerability Patterns Through Information Theory and Contrastive Learning
por: Nguyen, Van, et al.
Publicado: (2022)
por: Nguyen, Van, et al.
Publicado: (2022)
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
por: Hossain, Ismail, et al.
Publicado: (2026)
por: Hossain, Ismail, et al.
Publicado: (2026)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
por: Peng, Benji, et al.
Publicado: (2024)
por: Peng, Benji, et al.
Publicado: (2024)
Finetuning Large Language Models for Vulnerability Detection
por: Shestov, Alexey, et al.
Publicado: (2024)
por: Shestov, Alexey, et al.
Publicado: (2024)
Rethinking the Vulnerability of Concept Erasure and a New Method
por: Richardson, Alex D., et al.
Publicado: (2025)
por: Richardson, Alex D., et al.
Publicado: (2025)
Enhancing Vulnerability Reports with Automated and Augmented Description Summarization
por: Althebeiti, Hattan, et al.
Publicado: (2025)
por: Althebeiti, Hattan, et al.
Publicado: (2025)
Exploiting Efficiency Vulnerabilities in Dynamic Deep Learning Systems
por: Rathnasuriya, Ravishka, et al.
Publicado: (2025)
por: Rathnasuriya, Ravishka, et al.
Publicado: (2025)
ARVO: Atlas of Reproducible Vulnerabilities for Open Source Software
por: Mei, Xiang, et al.
Publicado: (2024)
por: Mei, Xiang, et al.
Publicado: (2024)
Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors
por: Sun, Luze, et al.
Publicado: (2026)
por: Sun, Luze, et al.
Publicado: (2026)
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
por: Li, Yuxi, et al.
Publicado: (2024)
por: Li, Yuxi, et al.
Publicado: (2024)
Training on Fake Labels: Mitigating Label Leakage in Split Learning via Secure Dimension Transformation
por: Jiang, Yukun, et al.
Publicado: (2024)
por: Jiang, Yukun, et al.
Publicado: (2024)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
por: Nurlanov, Zhakshylyk, et al.
Publicado: (2026)
por: Nurlanov, Zhakshylyk, et al.
Publicado: (2026)
Unveiling Privacy, Memorization, and Input Curvature Links
por: Ravikumar, Deepak, et al.
Publicado: (2024)
por: Ravikumar, Deepak, et al.
Publicado: (2024)
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
por: Jaffal, Niveen O., et al.
Publicado: (2025)
por: Jaffal, Niveen O., et al.
Publicado: (2025)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
por: Yue, Murong, et al.
Publicado: (2025)
por: Yue, Murong, et al.
Publicado: (2025)
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models
por: Krishna, Arjun, et al.
Publicado: (2025)
por: Krishna, Arjun, et al.
Publicado: (2025)
Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights
por: Fang, Xingli, et al.
Publicado: (2026)
por: Fang, Xingli, et al.
Publicado: (2026)
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
por: Abdali, Sara, et al.
Publicado: (2024)
por: Abdali, Sara, et al.
Publicado: (2024)
Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
por: Foroughi, Mohammad Hadi, et al.
Publicado: (2026)
por: Foroughi, Mohammad Hadi, et al.
Publicado: (2026)
Can Neural Decompilation Assist Vulnerability Prediction on Binary Code?
por: Cotroneo, D., et al.
Publicado: (2024)
por: Cotroneo, D., et al.
Publicado: (2024)
An Unbiased Transformer Source Code Learning with Semantic Vulnerability Graph
por: Islam, Nafis Tanveer, et al.
Publicado: (2023)
por: Islam, Nafis Tanveer, et al.
Publicado: (2023)
Your Privacy Depends on Others: Collusion Vulnerabilities in Individual Differential Privacy
por: Kaiser, Johannes, et al.
Publicado: (2026)
por: Kaiser, Johannes, et al.
Publicado: (2026)
Soft-Label Integration for Robust Toxicity Classification
por: Cheng, Zelei, et al.
Publicado: (2024)
por: Cheng, Zelei, et al.
Publicado: (2024)
Retrieval Augmented Anomaly Detection (RAAD): Nimble Model Adjustment Without Retraining
por: Pastoriza, Sam, et al.
Publicado: (2025)
por: Pastoriza, Sam, et al.
Publicado: (2025)
Disrupting Model Merging: A Parameter-Level Defense Without Sacrificing Accuracy
por: Junhao, Wei, et al.
Publicado: (2025)
por: Junhao, Wei, et al.
Publicado: (2025)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
por: Jiang, Fengqing, et al.
Publicado: (2024)
por: Jiang, Fengqing, et al.
Publicado: (2024)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
por: Wang, Zhun, et al.
Publicado: (2026)
por: Wang, Zhun, et al.
Publicado: (2026)
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
por: Liu, Fuqiang, et al.
Publicado: (2024)
por: Liu, Fuqiang, et al.
Publicado: (2024)
A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality
por: Huang, Hanbo, et al.
Publicado: (2024)
por: Huang, Hanbo, et al.
Publicado: (2024)
Relationship between Uncertainty in DNNs and Adversarial Attacks
por: Ogonna, Mabel, et al.
Publicado: (2024)
por: Ogonna, Mabel, et al.
Publicado: (2024)
Why Safety Probes Catch Liars But Miss Fanatics
por: Haralambiev, Kristiyan
Publicado: (2026)
por: Haralambiev, Kristiyan
Publicado: (2026)
AED: Automatic Discovery of Effective and Diverse Vulnerabilities for Autonomous Driving Policy with Large Language Models
por: Qiu, Le, et al.
Publicado: (2025)
por: Qiu, Le, et al.
Publicado: (2025)
Adversarial Vulnerability Under Temporal Concept Drift: A Longitudinal Study of Android Malware Detection
por: Sabbah, Ahmed, et al.
Publicado: (2026)
por: Sabbah, Ahmed, et al.
Publicado: (2026)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
por: Pal, Soumyadeep, et al.
Publicado: (2024)
por: Pal, Soumyadeep, et al.
Publicado: (2024)
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
por: Zhao, Mengnan, et al.
Publicado: (2026)
por: Zhao, Mengnan, et al.
Publicado: (2026)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
por: Chen, Depeng, et al.
Publicado: (2024)
por: Chen, Depeng, et al.
Publicado: (2024)
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
por: Lai, Zhenglin, et al.
Publicado: (2025)
por: Lai, Zhenglin, et al.
Publicado: (2025)
How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System
por: Zuo, Kaiwen, et al.
Publicado: (2025)
por: Zuo, Kaiwen, et al.
Publicado: (2025)
Ejemplares similares
-
BugSweeper: Function-Level Detection of Smart Contract Vulnerabilities Using Graph Neural Networks
por: Lee, Uisang, et al.
Publicado: (2025) -
Unveiling the Threat of Fraud Gangs to Graph Neural Networks: Multi-Target Graph Injection Attacks Against GNN-Based Fraud Detectors
por: Choi, Jinhyeok, et al.
Publicado: (2024) -
Statement-Level Vulnerability Detection: Learning Vulnerability Patterns Through Information Theory and Contrastive Learning
por: Nguyen, Van, et al.
Publicado: (2022) -
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
por: Hossain, Ismail, et al.
Publicado: (2026) -
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
por: Peng, Benji, et al.
Publicado: (2024)