Automated Consistency Analysis of LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patwardhan, Aditya, Vaidya, Vivek, Kundu, Ashish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Good LLM-Generated Password Policies Are?
von: Vaidya, Vivek, et al.
Veröffentlicht: (2025)
von: Vaidya, Vivek, et al.
Veröffentlicht: (2025)
GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs
von: Das, Sanjay, et al.
Veröffentlicht: (2024)
von: Das, Sanjay, et al.
Veröffentlicht: (2024)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
von: Lu, Ning, et al.
Veröffentlicht: (2025)
von: Lu, Ning, et al.
Veröffentlicht: (2025)
EVMbench: Evaluating AI Agents on Smart Contract Security
von: Wang, Justin, et al.
Veröffentlicht: (2026)
von: Wang, Justin, et al.
Veröffentlicht: (2026)
Linearizing Models for Efficient yet Robust Private Inference
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
Automated and Explainable Denial of Service Analysis for AI-Driven Intrusion Detection Systems
von: Yakubu, Paul Badu, et al.
Veröffentlicht: (2025)
von: Yakubu, Paul Badu, et al.
Veröffentlicht: (2025)
Modeling Behavioral Preferences of Cyber Adversaries Using Inverse Reinforcement Learning
von: Shinde, Aditya, et al.
Veröffentlicht: (2025)
von: Shinde, Aditya, et al.
Veröffentlicht: (2025)
On the Consistency of GNN Explanations for Malware Detection
von: Shokouhinejad, Hossein, et al.
Veröffentlicht: (2025)
von: Shokouhinejad, Hossein, et al.
Veröffentlicht: (2025)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
von: Wang, Zi, et al.
Veröffentlicht: (2024)
von: Wang, Zi, et al.
Veröffentlicht: (2024)
Generative Models are Self-Watermarked: Declaring Model Authentication through Re-Generation
von: Desu, Aditya, et al.
Veröffentlicht: (2024)
von: Desu, Aditya, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models for Phishing Detection, Self-Consistency, Faithfulness, and Explainability
von: Kuikel, Shova, et al.
Veröffentlicht: (2025)
von: Kuikel, Shova, et al.
Veröffentlicht: (2025)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2024)
von: Pal, Soumyadeep, et al.
Veröffentlicht: (2024)
Enhancing Vulnerability Reports with Automated and Augmented Description Summarization
von: Althebeiti, Hattan, et al.
Veröffentlicht: (2025)
von: Althebeiti, Hattan, et al.
Veröffentlicht: (2025)
Large-scale online deanonymization with LLMs
von: Lermen, Simon, et al.
Veröffentlicht: (2026)
von: Lermen, Simon, et al.
Veröffentlicht: (2026)
Scaling Trends for Data Poisoning in LLMs
von: Bowen, Dillon, et al.
Veröffentlicht: (2024)
von: Bowen, Dillon, et al.
Veröffentlicht: (2024)
Hybrid Temporal Differential Consistency Autoencoder for Efficient and Sustainable Anomaly Detection in Cyber-Physical Systems
von: Somma, Michael
Veröffentlicht: (2025)
von: Somma, Michael
Veröffentlicht: (2025)
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2026)
von: Choudhary, Sarthak, et al.
Veröffentlicht: (2026)
Towards Trustworthy AI: Secure Deepfake Detection using CNNs and Zero-Knowledge Proofs
von: Islam, H M Mohaimanul, et al.
Veröffentlicht: (2025)
von: Islam, H M Mohaimanul, et al.
Veröffentlicht: (2025)
CADRE: Customizable Assurance of Data Readiness in Privacy-Preserving Federated Learning
von: Hiniduma, Kaveen, et al.
Veröffentlicht: (2025)
von: Hiniduma, Kaveen, et al.
Veröffentlicht: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
Leveraging RAG for Training-Free Alignment of LLMs
von: Halloran, John T.
Veröffentlicht: (2026)
von: Halloran, John T.
Veröffentlicht: (2026)
AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema
von: Liu, Ting-Chun, et al.
Veröffentlicht: (2025)
von: Liu, Ting-Chun, et al.
Veröffentlicht: (2025)
Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models
von: Biskupski, Tom, et al.
Veröffentlicht: (2026)
von: Biskupski, Tom, et al.
Veröffentlicht: (2026)
RL-Finetuned LLMs for Privacy-Preserving Synthetic Rewriting
von: Shi, Zhan, et al.
Veröffentlicht: (2025)
von: Shi, Zhan, et al.
Veröffentlicht: (2025)
Agentic Misalignment: How LLMs Could Be Insider Threats
von: Lynch, Aengus, et al.
Veröffentlicht: (2025)
von: Lynch, Aengus, et al.
Veröffentlicht: (2025)
Fast Exact Unlearning for In-Context Learning Data for LLMs
von: Muresanu, Andrei I., et al.
Veröffentlicht: (2024)
von: Muresanu, Andrei I., et al.
Veröffentlicht: (2024)
Attention Tracker: Detecting Prompt Injection Attacks in LLMs
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
von: Hung, Kuo-Han, et al.
Veröffentlicht: (2024)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
von: Wang, Erchi, et al.
Veröffentlicht: (2026)
von: Wang, Erchi, et al.
Veröffentlicht: (2026)
Permissioned LLMs: Enforcing Access Control in Large Language Models
von: Jayaraman, Bargav, et al.
Veröffentlicht: (2025)
von: Jayaraman, Bargav, et al.
Veröffentlicht: (2025)
XBreaking: Understanding how LLMs security alignment can be broken
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
von: Arazzi, Marco, et al.
Veröffentlicht: (2025)
LAMD: Context-driven Android Malware Detection and Classification with LLMs
von: Qian, Xingzhi, et al.
Veröffentlicht: (2025)
von: Qian, Xingzhi, et al.
Veröffentlicht: (2025)
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
von: Park, Sangwoo, et al.
Veröffentlicht: (2026)
von: Park, Sangwoo, et al.
Veröffentlicht: (2026)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
von: Nurlanov, Zhakshylyk, et al.
Veröffentlicht: (2026)
Automated Classification of Cybercrime Complaints using Transformer-based Language Models for Hinglish Texts
von: Rani, Nanda, et al.
Veröffentlicht: (2024)
von: Rani, Nanda, et al.
Veröffentlicht: (2024)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2025)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
TracLLM: A Generic Framework for Attributing Long Context LLMs
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
von: Wang, Yanting, et al.
Veröffentlicht: (2025)
POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
von: Li, Xinyu, et al.
Veröffentlicht: (2025)
von: Li, Xinyu, et al.
Veröffentlicht: (2025)
PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Good LLM-Generated Password Policies Are?
von: Vaidya, Vivek, et al.
Veröffentlicht: (2025) -
GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs
von: Das, Sanjay, et al.
Veröffentlicht: (2024) -
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
von: Lu, Ning, et al.
Veröffentlicht: (2025) -
EVMbench: Evaluating AI Agents on Smart Contract Security
von: Wang, Justin, et al.
Veröffentlicht: (2026) -
Linearizing Models for Efficient yet Robust Private Inference
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)