Guardado en:
| Autores principales: | Patwardhan, Aditya, Vaidya, Vivek, Kundu, Ashish |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.07036 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Good LLM-Generated Password Policies Are?
por: Vaidya, Vivek, et al.
Publicado: (2025)
por: Vaidya, Vivek, et al.
Publicado: (2025)
GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs
por: Das, Sanjay, et al.
Publicado: (2024)
por: Das, Sanjay, et al.
Publicado: (2024)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
por: Lu, Ning, et al.
Publicado: (2025)
por: Lu, Ning, et al.
Publicado: (2025)
EVMbench: Evaluating AI Agents on Smart Contract Security
por: Wang, Justin, et al.
Publicado: (2026)
por: Wang, Justin, et al.
Publicado: (2026)
Linearizing Models for Efficient yet Robust Private Inference
por: Sarkar, Sreetama, et al.
Publicado: (2024)
por: Sarkar, Sreetama, et al.
Publicado: (2024)
Modeling Behavioral Preferences of Cyber Adversaries Using Inverse Reinforcement Learning
por: Shinde, Aditya, et al.
Publicado: (2025)
por: Shinde, Aditya, et al.
Publicado: (2025)
Automated and Explainable Denial of Service Analysis for AI-Driven Intrusion Detection Systems
por: Yakubu, Paul Badu, et al.
Publicado: (2025)
por: Yakubu, Paul Badu, et al.
Publicado: (2025)
On the Consistency of GNN Explanations for Malware Detection
por: Shokouhinejad, Hossein, et al.
Publicado: (2025)
por: Shokouhinejad, Hossein, et al.
Publicado: (2025)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
por: Wang, Zi, et al.
Publicado: (2024)
por: Wang, Zi, et al.
Publicado: (2024)
Generative Models are Self-Watermarked: Declaring Model Authentication through Re-Generation
por: Desu, Aditya, et al.
Publicado: (2024)
por: Desu, Aditya, et al.
Publicado: (2024)
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
por: Choudhary, Sarthak, et al.
Publicado: (2026)
por: Choudhary, Sarthak, et al.
Publicado: (2026)
Evaluating Large Language Models for Phishing Detection, Self-Consistency, Faithfulness, and Explainability
por: Kuikel, Shova, et al.
Publicado: (2025)
por: Kuikel, Shova, et al.
Publicado: (2025)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
por: Pal, Soumyadeep, et al.
Publicado: (2024)
por: Pal, Soumyadeep, et al.
Publicado: (2024)
Towards Trustworthy AI: Secure Deepfake Detection using CNNs and Zero-Knowledge Proofs
por: Islam, H M Mohaimanul, et al.
Publicado: (2025)
por: Islam, H M Mohaimanul, et al.
Publicado: (2025)
CADRE: Customizable Assurance of Data Readiness in Privacy-Preserving Federated Learning
por: Hiniduma, Kaveen, et al.
Publicado: (2025)
por: Hiniduma, Kaveen, et al.
Publicado: (2025)
Enhancing Vulnerability Reports with Automated and Augmented Description Summarization
por: Althebeiti, Hattan, et al.
Publicado: (2025)
por: Althebeiti, Hattan, et al.
Publicado: (2025)
Large-scale online deanonymization with LLMs
por: Lermen, Simon, et al.
Publicado: (2026)
por: Lermen, Simon, et al.
Publicado: (2026)
Scaling Trends for Data Poisoning in LLMs
por: Bowen, Dillon, et al.
Publicado: (2024)
por: Bowen, Dillon, et al.
Publicado: (2024)
Hybrid Temporal Differential Consistency Autoencoder for Efficient and Sustainable Anomaly Detection in Cyber-Physical Systems
por: Somma, Michael
Publicado: (2025)
por: Somma, Michael
Publicado: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
por: Chen, Zhuowei, et al.
Publicado: (2025)
por: Chen, Zhuowei, et al.
Publicado: (2025)
Leveraging RAG for Training-Free Alignment of LLMs
por: Halloran, John T.
Publicado: (2026)
por: Halloran, John T.
Publicado: (2026)
AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema
por: Liu, Ting-Chun, et al.
Publicado: (2025)
por: Liu, Ting-Chun, et al.
Publicado: (2025)
Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models
por: Biskupski, Tom, et al.
Publicado: (2026)
por: Biskupski, Tom, et al.
Publicado: (2026)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
por: Shi, Zhan, et al.
Publicado: (2025)
por: Shi, Zhan, et al.
Publicado: (2025)
Agentic Misalignment: How LLMs Could Be Insider Threats
por: Lynch, Aengus, et al.
Publicado: (2025)
por: Lynch, Aengus, et al.
Publicado: (2025)
Fast Exact Unlearning for In-Context Learning Data for LLMs
por: Muresanu, Andrei I., et al.
Publicado: (2024)
por: Muresanu, Andrei I., et al.
Publicado: (2024)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
por: Hung, Kuo-Han, et al.
Publicado: (2024)
por: Hung, Kuo-Han, et al.
Publicado: (2024)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
por: Wang, Erchi, et al.
Publicado: (2026)
por: Wang, Erchi, et al.
Publicado: (2026)
Permissioned LLMs: Enforcing Access Control in Large Language Models
por: Jayaraman, Bargav, et al.
Publicado: (2025)
por: Jayaraman, Bargav, et al.
Publicado: (2025)
XBreaking: Understanding how LLMs security alignment can be broken
por: Arazzi, Marco, et al.
Publicado: (2025)
por: Arazzi, Marco, et al.
Publicado: (2025)
LAMD: Context-driven Android Malware Detection and Classification with LLMs
por: Qian, Xingzhi, et al.
Publicado: (2025)
por: Qian, Xingzhi, et al.
Publicado: (2025)
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
por: Park, Sangwoo, et al.
Publicado: (2026)
por: Park, Sangwoo, et al.
Publicado: (2026)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
por: Nurlanov, Zhakshylyk, et al.
Publicado: (2026)
por: Nurlanov, Zhakshylyk, et al.
Publicado: (2026)
Automated Classification of Cybercrime Complaints using Transformer-based Language Models for Hinglish Texts
por: Rani, Nanda, et al.
Publicado: (2024)
por: Rani, Nanda, et al.
Publicado: (2024)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
por: Panfilov, Alexander, et al.
Publicado: (2025)
por: Panfilov, Alexander, et al.
Publicado: (2025)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
por: Bazinska, Julia, et al.
Publicado: (2025)
por: Bazinska, Julia, et al.
Publicado: (2025)
TracLLM: A Generic Framework for Attributing Long Context LLMs
por: Wang, Yanting, et al.
Publicado: (2025)
por: Wang, Yanting, et al.
Publicado: (2025)
POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
por: Li, Xinyu, et al.
Publicado: (2025)
por: Li, Xinyu, et al.
Publicado: (2025)
PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs
por: Xue, Jiaqi, et al.
Publicado: (2025)
por: Xue, Jiaqi, et al.
Publicado: (2025)
Ejemplares similares
-
How Good LLM-Generated Password Policies Are?
por: Vaidya, Vivek, et al.
Publicado: (2025) -
GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs
por: Das, Sanjay, et al.
Publicado: (2024) -
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
por: Lu, Ning, et al.
Publicado: (2025) -
EVMbench: Evaluating AI Agents on Smart Contract Security
por: Wang, Justin, et al.
Publicado: (2026) -
Linearizing Models for Efficient yet Robust Private Inference
por: Sarkar, Sreetama, et al.
Publicado: (2024)