Toward Reliable, Safe, and Secure LLMs for Scientific Applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chaturvedi, Saket Sanjeev, Bergerson, Joshua, Mallick, Tanwi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
von: Chen, Beitao, et al.
Veröffentlicht: (2025)
von: Chen, Beitao, et al.
Veröffentlicht: (2025)
SafeText: Safe Text-to-image Models via Aligning the Text Encoder
von: Hu, Yuepeng, et al.
Veröffentlicht: (2025)
von: Hu, Yuepeng, et al.
Veröffentlicht: (2025)
Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2025)
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2025)
Towards an End-to-End (E2E) Adversarial Learning and Application in the Physical World
von: Biton, Dudi, et al.
Veröffentlicht: (2025)
von: Biton, Dudi, et al.
Veröffentlicht: (2025)
Landscape More Secure Than Portrait? Zooming Into the Directionality of Digital Images With Security Implications
von: Lorch, Benedikt, et al.
Veröffentlicht: (2024)
von: Lorch, Benedikt, et al.
Veröffentlicht: (2024)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Is Diffusion Model Safe? Severe Data Leakage via Gradient-Guided Diffusion Model
von: Meng, Jiayang, et al.
Veröffentlicht: (2024)
von: Meng, Jiayang, et al.
Veröffentlicht: (2024)
SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
von: Rong, Xuankun, et al.
Veröffentlicht: (2025)
von: Rong, Xuankun, et al.
Veröffentlicht: (2025)
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
von: Jia, Xiaojun, et al.
Veröffentlicht: (2025)
Securing Face and Fingerprint Templates in Humanitarian Biometric Systems
von: Stragapede, Giuseppe, et al.
Veröffentlicht: (2025)
von: Stragapede, Giuseppe, et al.
Veröffentlicht: (2025)
Accuracy Limits as a Barrier to Biometric System Security
von: Durbet, Axel, et al.
Veröffentlicht: (2024)
von: Durbet, Axel, et al.
Veröffentlicht: (2024)
SecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations
von: Ali, Mohammed Himayath, et al.
Veröffentlicht: (2026)
von: Ali, Mohammed Himayath, et al.
Veröffentlicht: (2026)
Secure Data Access in Cloud Environments Using Quantum Cryptography
von: Lakshmi, S. Vasavi Venkata, et al.
Veröffentlicht: (2025)
von: Lakshmi, S. Vasavi Venkata, et al.
Veröffentlicht: (2025)
Machine Learning Security against Data Poisoning: Are We There Yet?
von: Cinà, Antonio Emanuele, et al.
Veröffentlicht: (2022)
von: Cinà, Antonio Emanuele, et al.
Veröffentlicht: (2022)
Secure and Scalable Face Retrieval via Cancelable Product Quantization
von: Tang, Haomiao, et al.
Veröffentlicht: (2025)
von: Tang, Haomiao, et al.
Veröffentlicht: (2025)
Robust Provably Secure Image Steganography via Latent Iterative Optimization
von: Li, Yanan, et al.
Veröffentlicht: (2026)
von: Li, Yanan, et al.
Veröffentlicht: (2026)
Anomaly Unveiled: Securing Image Classification against Adversarial Patch Attacks
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2024)
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2024)
Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
von: Cao, Jie, et al.
Veröffentlicht: (2025)
von: Cao, Jie, et al.
Veröffentlicht: (2025)
Hiding Your Signals: A Security Analysis of PPG-based Biometric Authentication
von: Li, Lin, et al.
Veröffentlicht: (2022)
von: Li, Lin, et al.
Veröffentlicht: (2022)
AI-Driven Secure Data Sharing: A Trustworthy and Privacy-Preserving Approach
von: Amin, Al, et al.
Veröffentlicht: (2025)
von: Amin, Al, et al.
Veröffentlicht: (2025)
Secure Seed-Based Multi-bit Watermarking for Diffusion Models from First Principles
von: Gesny, Enoal, et al.
Veröffentlicht: (2026)
von: Gesny, Enoal, et al.
Veröffentlicht: (2026)
Lightweight True In-Pixel Encryption with FeFET Enabled Pixel Design for Secure Imaging
von: Udoy, Md Rahatul Islam, et al.
Veröffentlicht: (2026)
von: Udoy, Md Rahatul Islam, et al.
Veröffentlicht: (2026)
DLOVE: A new Security Evaluation Tool for Deep Learning Based Watermarking Techniques
von: Padhi, Sudev Kumar, et al.
Veröffentlicht: (2024)
von: Padhi, Sudev Kumar, et al.
Veröffentlicht: (2024)
SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts
von: Wu, Xiaodong, et al.
Veröffentlicht: (2025)
von: Wu, Xiaodong, et al.
Veröffentlicht: (2025)
Advancing Security in AI Systems: A Novel Approach to Detecting Backdoors in Deep Neural Networks
von: Hossain, Khondoker Murad, et al.
Veröffentlicht: (2024)
von: Hossain, Khondoker Murad, et al.
Veröffentlicht: (2024)
Training-Free Color-Aware Adversarial Diffusion Sanitization for Diffusion Stegomalware Defense at Security Gateways
von: Frants, Vladimir, et al.
Veröffentlicht: (2025)
von: Frants, Vladimir, et al.
Veröffentlicht: (2025)
Hypersphere Secure Sketch Revisited: Probabilistic Linear Regression Attack on IronMask in Multiple Usage
von: Zhu, Pengxu, et al.
Veröffentlicht: (2024)
von: Zhu, Pengxu, et al.
Veröffentlicht: (2024)
Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
von: Aerni, Michael, et al.
Veröffentlicht: (2025)
von: Aerni, Michael, et al.
Veröffentlicht: (2025)
OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs
von: Wang, Xin, et al.
Veröffentlicht: (2026)
von: Wang, Xin, et al.
Veröffentlicht: (2026)
AttackNet: Enhancing Biometric Security via Tailored Convolutional Neural Network Architectures for Liveness Detection
von: Kuznetsov, Oleksandr, et al.
Veröffentlicht: (2024)
von: Kuznetsov, Oleksandr, et al.
Veröffentlicht: (2024)
Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security
von: Fan, Yihe, et al.
Veröffentlicht: (2024)
von: Fan, Yihe, et al.
Veröffentlicht: (2024)
Leak and Learn: An Attacker's Cookbook to Train Using Leaked Data from Federated Learning
von: Zhao, Joshua C., et al.
Veröffentlicht: (2024)
von: Zhao, Joshua C., et al.
Veröffentlicht: (2024)
Towards Physical World Backdoor Attacks against Skeleton Action Recognition
von: Zheng, Qichen, et al.
Veröffentlicht: (2024)
von: Zheng, Qichen, et al.
Veröffentlicht: (2024)
DREW : Towards Robust Data Provenance by Leveraging Error-Controlled Watermarking
von: Saberi, Mehrdad, et al.
Veröffentlicht: (2024)
von: Saberi, Mehrdad, et al.
Veröffentlicht: (2024)
A Machine Learning-Based Secure Face Verification Scheme and Its Applications to Digital Surveillance
von: Wang, Huan-Chih, et al.
Veröffentlicht: (2024)
von: Wang, Huan-Chih, et al.
Veröffentlicht: (2024)
Towards Physically Realizable Adversarial Attenuation Patch against SAR Object Detection
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
$\mathbf{S^2LM}$: Towards Semantic Steganography via Large Language Models
von: Wu, Huanqi, et al.
Veröffentlicht: (2025)
von: Wu, Huanqi, et al.
Veröffentlicht: (2025)
Nearest is Not Dearest: Towards Practical Defense against Quantization-conditioned Backdoor Attacks
von: Li, Boheng, et al.
Veröffentlicht: (2024)
von: Li, Boheng, et al.
Veröffentlicht: (2024)
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
FoolSDEdit: Deceptively Steering Your Edits Towards Targeted Attribute-aware Distribution
von: Zhou, Qi, et al.
Veröffentlicht: (2024)
von: Zhou, Qi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
von: Chen, Beitao, et al.
Veröffentlicht: (2025) -
SafeText: Safe Text-to-image Models via Aligning the Text Encoder
von: Hu, Yuepeng, et al.
Veröffentlicht: (2025) -
Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application
von: Xu, Wenzhuo, et al.
Veröffentlicht: (2025) -
Towards an End-to-End (E2E) Adversarial Learning and Application in the Physical World
von: Biton, Dudi, et al.
Veröffentlicht: (2025) -
Landscape More Secure Than Portrait? Zooming Into the Directionality of Digital Images With Security Implications
von: Lorch, Benedikt, et al.
Veröffentlicht: (2024)