BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gan, Jialing, Dong, Junhao, Li, Songze |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
von: Jawad, Huseein, et al.
Veröffentlicht: (2025)
von: Jawad, Huseein, et al.
Veröffentlicht: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
von: Freenor, Michael, et al.
Veröffentlicht: (2025)
von: Freenor, Michael, et al.
Veröffentlicht: (2025)
DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
von: Sun, Xiongtao, et al.
Veröffentlicht: (2024)
von: Sun, Xiongtao, et al.
Veröffentlicht: (2024)
Auditing Prompt Caching in Language Model APIs
von: Gu, Chenchen, et al.
Veröffentlicht: (2025)
von: Gu, Chenchen, et al.
Veröffentlicht: (2025)
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
von: Kabir, Md Rysul, et al.
Veröffentlicht: (2026)
von: Kabir, Md Rysul, et al.
Veröffentlicht: (2026)
Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks
von: Iyer, Karthik Raghu, et al.
Veröffentlicht: (2026)
von: Iyer, Karthik Raghu, et al.
Veröffentlicht: (2026)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
von: Li, Caihua, et al.
Veröffentlicht: (2024)
von: Li, Caihua, et al.
Veröffentlicht: (2024)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
von: Wang, Peiran, et al.
Veröffentlicht: (2026)
von: Wang, Peiran, et al.
Veröffentlicht: (2026)
Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
ContextLeak: Auditing Leakage in Private In-Context Learning Methods
von: Choi, Jacob, et al.
Veröffentlicht: (2025)
von: Choi, Jacob, et al.
Veröffentlicht: (2025)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
von: Shen, Guobin, et al.
Veröffentlicht: (2025)
Hidden Data Privacy Breaches in Federated Learning
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
von: Gong, Xueluan, et al.
Veröffentlicht: (2024)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
von: Li, Yucheng, et al.
Veröffentlicht: (2025)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
von: Das, Saswat, et al.
Veröffentlicht: (2025)
von: Das, Saswat, et al.
Veröffentlicht: (2025)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
von: An, Hengyu, et al.
Veröffentlicht: (2025)
von: An, Hengyu, et al.
Veröffentlicht: (2025)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
von: Xin, Yuan, et al.
Veröffentlicht: (2026)
von: Xin, Yuan, et al.
Veröffentlicht: (2026)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
von: Yang, Yong, et al.
Veröffentlicht: (2024)
von: Yang, Yong, et al.
Veröffentlicht: (2024)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
von: Singh, Himanshu, et al.
Veröffentlicht: (2026)
von: Singh, Himanshu, et al.
Veröffentlicht: (2026)
Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
von: Yan, Lecheng, et al.
Veröffentlicht: (2026)
SecureLLM: Using Compositionality to Build Provably Secure Language Models for Private, Sensitive, and Secret Data
von: Alabdulkareem, Abdulrahman, et al.
Veröffentlicht: (2024)
von: Alabdulkareem, Abdulrahman, et al.
Veröffentlicht: (2024)
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2025)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework
von: Li, Junhao, et al.
Veröffentlicht: (2025)
von: Li, Junhao, et al.
Veröffentlicht: (2025)
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models
von: Liang, Zi, et al.
Veröffentlicht: (2024)
von: Liang, Zi, et al.
Veröffentlicht: (2024)
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
von: Wang, Zhuoshang, et al.
Veröffentlicht: (2026)
von: Wang, Zhuoshang, et al.
Veröffentlicht: (2026)
FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework
von: Béjar, Mario Rodríguez, et al.
Veröffentlicht: (2026)
von: Béjar, Mario Rodríguez, et al.
Veröffentlicht: (2026)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
von: Cai, Will, et al.
Veröffentlicht: (2025)
von: Cai, Will, et al.
Veröffentlicht: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
von: Pathade, Chetan
Veröffentlicht: (2025)
von: Pathade, Chetan
Veröffentlicht: (2025)
Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals
von: Akinrele, Akindoyin, et al.
Veröffentlicht: (2026)
von: Akinrele, Akindoyin, et al.
Veröffentlicht: (2026)
Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning
von: Chen, Chaoran, et al.
Veröffentlicht: (2026)
von: Chen, Chaoran, et al.
Veröffentlicht: (2026)
Is the Digital Forensics and Incident Response Pipeline Ready for Text-Based Threats in LLM Era?
von: Bhandarkar, Avanti, et al.
Veröffentlicht: (2024)
von: Bhandarkar, Avanti, et al.
Veröffentlicht: (2024)
Fingerprinting LLMs via Prompt Injection
von: Hu, Yuepeng, et al.
Veröffentlicht: (2025)
von: Hu, Yuepeng, et al.
Veröffentlicht: (2025)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
von: Hastuti, Rochana Prih, et al.
Veröffentlicht: (2025)
Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method
von: Zhang, Weichao, et al.
Veröffentlicht: (2024)
von: Zhang, Weichao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
von: Jawad, Huseein, et al.
Veröffentlicht: (2025) -
Prompt Optimization and Evaluation for LLM Automated Red Teaming
von: Freenor, Michael, et al.
Veröffentlicht: (2025) -
DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
von: Sun, Xiongtao, et al.
Veröffentlicht: (2024) -
Auditing Prompt Caching in Language Model APIs
von: Gu, Chenchen, et al.
Veröffentlicht: (2025) -
Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks
von: Kabir, Md Rysul, et al.
Veröffentlicht: (2026)