Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Leonesi, Matteo, Belardinelli, Francesco, Corradini, Flavio, Piangerelli, Marco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
von: Bercovich, Ivan, et al.
Veröffentlicht: (2026)
von: Bercovich, Ivan, et al.
Veröffentlicht: (2026)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
von: Rosenblatt, Lucas, et al.
Veröffentlicht: (2026)
von: Rosenblatt, Lucas, et al.
Veröffentlicht: (2026)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
von: Blanco-Justicia, Alberto, et al.
Veröffentlicht: (2024)
von: Blanco-Justicia, Alberto, et al.
Veröffentlicht: (2024)
SALLIE: Safeguarding Against Latent Language & Image Exploits
von: Azov, Guy, et al.
Veröffentlicht: (2026)
von: Azov, Guy, et al.
Veröffentlicht: (2026)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
von: Seddik, Issam, et al.
Veröffentlicht: (2025)
von: Seddik, Issam, et al.
Veröffentlicht: (2025)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
von: Young, Richard J., et al.
Veröffentlicht: (2026)
von: Young, Richard J., et al.
Veröffentlicht: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
Lightweight LLMs for Network Attack Detection in IoT Networks
von: Sudasinghe, Piyumi Bhagya, et al.
Veröffentlicht: (2026)
von: Sudasinghe, Piyumi Bhagya, et al.
Veröffentlicht: (2026)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
von: Lelle, Travis
Veröffentlicht: (2026)
von: Lelle, Travis
Veröffentlicht: (2026)
Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
von: Beltoft, Stine, et al.
Veröffentlicht: (2025)
von: Beltoft, Stine, et al.
Veröffentlicht: (2025)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
Scalable and Verifiable Federated Learning for Cross-Institution Financial Fraud Detection
von: Panth, Prajwal, et al.
Veröffentlicht: (2026)
von: Panth, Prajwal, et al.
Veröffentlicht: (2026)
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
von: Gu, Yongtong, et al.
Veröffentlicht: (2026)
von: Gu, Yongtong, et al.
Veröffentlicht: (2026)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
von: Morales, Jaime, et al.
Veröffentlicht: (2026)
von: Morales, Jaime, et al.
Veröffentlicht: (2026)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
von: Lucas, Tom, et al.
Veröffentlicht: (2026)
von: Lucas, Tom, et al.
Veröffentlicht: (2026)
Unlearning at Scale: Implementing the Right to be Forgotten in Large Language Models
von: X, Abdullah
Veröffentlicht: (2025)
von: X, Abdullah
Veröffentlicht: (2025)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
von: Othman, Refat
Veröffentlicht: (2026)
von: Othman, Refat
Veröffentlicht: (2026)
Retrieval Augmented Classification for Confidential Documents
von: Chang, Yeseul E., et al.
Veröffentlicht: (2026)
von: Chang, Yeseul E., et al.
Veröffentlicht: (2026)
Whisper Leak: a side-channel attack on Large Language Models
von: McDonald, Geoff, et al.
Veröffentlicht: (2025)
von: McDonald, Geoff, et al.
Veröffentlicht: (2025)
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
Do Schwartz Higher-Order Values Help Sentence-Level Human Value Detection? A Study of Hierarchical Gating and Calibration
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
Sensitivity Uncertainty Alignment in Large Language Models
von: Hiremath, Prakul Sunil, et al.
Veröffentlicht: (2026)
von: Hiremath, Prakul Sunil, et al.
Veröffentlicht: (2026)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
von: Gowda, Ishrith
Veröffentlicht: (2026)
von: Gowda, Ishrith
Veröffentlicht: (2026)
Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management
von: Arora, Sunil, et al.
Veröffentlicht: (2025)
von: Arora, Sunil, et al.
Veröffentlicht: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
von: Chona, Alankrit, et al.
Veröffentlicht: (2026)
von: Chona, Alankrit, et al.
Veröffentlicht: (2026)
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction
von: Garcia, Gabriel
Veröffentlicht: (2026)
von: Garcia, Gabriel
Veröffentlicht: (2026)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
von: Nathanson, Samuel, et al.
Veröffentlicht: (2025)
von: Nathanson, Samuel, et al.
Veröffentlicht: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
von: Gaikwad, Madhava
Veröffentlicht: (2025)
von: Gaikwad, Madhava
Veröffentlicht: (2025)
How Worrying Are Privacy Attacks Against Machine Learning?
von: Domingo-Ferrer, Josep
Veröffentlicht: (2025)
von: Domingo-Ferrer, Josep
Veröffentlicht: (2025)
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
von: Bommarito II, Michael J.
Veröffentlicht: (2025)
von: Bommarito II, Michael J.
Veröffentlicht: (2025)
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
von: Chen, Renmiao, et al.
Veröffentlicht: (2025)
von: Chen, Renmiao, et al.
Veröffentlicht: (2025)
Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale
von: Tsai, Elisa, et al.
Veröffentlicht: (2025)
von: Tsai, Elisa, et al.
Veröffentlicht: (2025)
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
von: Li, Shenghao
Veröffentlicht: (2025)
von: Li, Shenghao
Veröffentlicht: (2025)
Attacking interpretable NLP systems
von: Abdukhamidov, Eldor, et al.
Veröffentlicht: (2025)
von: Abdukhamidov, Eldor, et al.
Veröffentlicht: (2025)
KidsNanny: A Two-Stage Multimodal Content Moderation Pipeline Integrating Visual Classification, Object Detection, OCR, and Contextual Reasoning for Child Safety
von: Panchal, Viraj, et al.
Veröffentlicht: (2026)
von: Panchal, Viraj, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
von: DeLeeuw, Caleb
Veröffentlicht: (2026) -
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
von: Bercovich, Ivan, et al.
Veröffentlicht: (2026) -
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
von: Rosenblatt, Lucas, et al.
Veröffentlicht: (2026) -
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
von: Blanco-Justicia, Alberto, et al.
Veröffentlicht: (2024) -
SALLIE: Safeguarding Against Latent Language & Image Exploits
von: Azov, Guy, et al.
Veröffentlicht: (2026)