Benchmarking Deception Probes via Black-to-White Performance Boosts
Fuente:
arXiv
Saved in:
| Main Authors: | Parrack, Avi, Attubato, Carlo Leonardo, Heimersheim, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
by: Yagoubi, Faouzi El, et al.
Published: (2026)
by: Yagoubi, Faouzi El, et al.
Published: (2026)
Computable Gap Assessment of Artificial Intelligence Governance in Children's Centres: Evidence-Mechanism-Governance-Indicator Modelling of UNICEF's Guidance on AI and Children 3.0 Based on the Graph-GAP Framework
by: Meng, Wei
Published: (2025)
by: Meng, Wei
Published: (2025)
AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
AI Ethics Principles in Practice: Perspectives of Designers and Developers
by: Sanderson, Conrad, et al.
Published: (2021)
by: Sanderson, Conrad, et al.
Published: (2021)
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026)
by: Wang, Yanshu, et al.
Published: (2026)
ChatGPT Based Data Augmentation for Improved Parameter-Efficient Debiasing of LLMs
by: Han, Pengrui, et al.
Published: (2024)
by: Han, Pengrui, et al.
Published: (2024)
MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
by: Dékány, Csaba, et al.
Published: (2025)
by: Dékány, Csaba, et al.
Published: (2025)
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
by: Schappacher-Tilp, Gudrun, et al.
Published: (2026)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
Recent Advances in Data-Driven Business Process Management
by: Ackermann, Lars, et al.
Published: (2024)
by: Ackermann, Lars, et al.
Published: (2024)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
by: Hossain, Ariyan, et al.
Published: (2025)
by: Hossain, Ariyan, et al.
Published: (2025)
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
by: Bahar, Atmane Ayoub Mansour, et al.
Published: (2024)
Revealing Hidden Bias in AI: Lessons from Large Language Models
by: Beatty, Django, et al.
Published: (2024)
by: Beatty, Django, et al.
Published: (2024)
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
by: Zhu, Pengyun, et al.
Published: (2026)
by: Zhu, Pengyun, et al.
Published: (2026)
The AI Fiction Paradox
by: Elkins, Katherine
Published: (2026)
by: Elkins, Katherine
Published: (2026)
AI-Powered Citation Auditing: A Zero-Assumption Protocol for Systematic Reference Verification in Academic Research
by: van Rensburg, L. J. Janse
Published: (2025)
by: van Rensburg, L. J. Janse
Published: (2025)
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
by: Gautam, Sushant, et al.
Published: (2026)
by: Gautam, Sushant, et al.
Published: (2026)
AVEC: Bootstrapping Privacy for Local LLMs
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
by: Young, Richard
Published: (2025)
by: Young, Richard
Published: (2025)
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
by: Hartmann, David, et al.
Published: (2026)
by: Hartmann, David, et al.
Published: (2026)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
by: Stewart, Ian, et al.
Published: (2024)
by: Stewart, Ian, et al.
Published: (2024)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
by: Cohen, Liran, et al.
Published: (2025)
by: Cohen, Liran, et al.
Published: (2025)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
by: Zmanovskii, Nikita
Published: (2025)
by: Zmanovskii, Nikita
Published: (2025)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
by: Leonesi, Matteo, et al.
Published: (2026)
by: Leonesi, Matteo, et al.
Published: (2026)
ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code
by: Madan, Kapil
Published: (2025)
by: Madan, Kapil
Published: (2025)
AI to Learn 2.0: A Deliverable-Oriented Governance Framework and Maturity Rubric for Opaque AI in Learning-Intensive Domains
by: Shintani, Seine A.
Published: (2026)
by: Shintani, Seine A.
Published: (2026)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
The Ethics Engine: A Modular Pipeline for Accessible Psychometric Assessment of Large Language Models
by: Van Clief, Jake, et al.
Published: (2025)
by: Van Clief, Jake, et al.
Published: (2025)
Synthetic emotions and consciousness: exploring architectural boundaries
by: Borotschnig, Hermann
Published: (2025)
by: Borotschnig, Hermann
Published: (2025)
Cultural Encoding in Large Language Models: The Existence Gap in AI-Mediated Brand Discovery
by: Junyao, Huang, et al.
Published: (2025)
by: Junyao, Huang, et al.
Published: (2025)
Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
by: Kryshtal, Andrii
Published: (2026)
by: Kryshtal, Andrii
Published: (2026)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
by: Lucas, Tom, et al.
Published: (2026)
by: Lucas, Tom, et al.
Published: (2026)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
by: Beltoft, Stine, et al.
Published: (2025)
by: Beltoft, Stine, et al.
Published: (2025)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
by: Wang, Xinyue, et al.
Published: (2026)
by: Wang, Xinyue, et al.
Published: (2026)
Resolving Ethics Trade-offs in Implementing Responsible AI
by: Sanderson, Conrad, et al.
Published: (2024)
by: Sanderson, Conrad, et al.
Published: (2024)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source Models
by: Shashidhar, Sumuk, et al.
Published: (2023)
by: Shashidhar, Sumuk, et al.
Published: (2023)
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values
by: Chun, Jon, et al.
Published: (2024)
by: Chun, Jon, et al.
Published: (2024)
Implicit Geographic Inference in LLM Medical Triage: Language-Driven Disparities in Emergency Recommendations
by: Wong, Qi Han
Published: (2026)
by: Wong, Qi Han
Published: (2026)
Similar Items
-
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
by: Yagoubi, Faouzi El, et al.
Published: (2026) -
Computable Gap Assessment of Artificial Intelligence Governance in Children's Centres: Evidence-Mechanism-Governance-Indicator Modelling of UNICEF's Guidance on AI and Children 3.0 Based on the Graph-GAP Framework
by: Meng, Wei
Published: (2025) -
AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
by: Gaikwad, Madhava
Published: (2025) -
AI Ethics Principles in Practice: Perspectives of Designers and Developers
by: Sanderson, Conrad, et al.
Published: (2021) -
How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks
by: Wang, Yanshu, et al.
Published: (2026)