Mechanistic Interpretability in the Presence of Architectural Obfuscation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Florencio, Marcos, Barton, Thomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An End-to-End Homomorphically Encrypted Neural Network
von: Florencio, Marcos, et al.
Veröffentlicht: (2025)
von: Florencio, Marcos, et al.
Veröffentlicht: (2025)
Obfuscated Memory Malware Detection
von: P, Sharmila S, et al.
Veröffentlicht: (2024)
von: P, Sharmila S, et al.
Veröffentlicht: (2024)
Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
von: Wang, Ziwei, et al.
Veröffentlicht: (2026)
von: Wang, Ziwei, et al.
Veröffentlicht: (2026)
Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice
von: Zhang, Gehao, et al.
Veröffentlicht: (2025)
von: Zhang, Gehao, et al.
Veröffentlicht: (2025)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025)
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025)
FedAdOb: Privacy-Preserving Federated Deep Learning with Adaptive Obfuscation
von: Gu, Hanlin, et al.
Veröffentlicht: (2024)
von: Gu, Hanlin, et al.
Veröffentlicht: (2024)
Towards Privacy-Preserving LLM Inference via Covariant Obfuscation (Technical Report)
von: Lin, Yu, et al.
Veröffentlicht: (2026)
von: Lin, Yu, et al.
Veröffentlicht: (2026)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
von: Mohseni, Seyedreza, et al.
Veröffentlicht: (2024)
von: Mohseni, Seyedreza, et al.
Veröffentlicht: (2024)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
von: Zhang, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Kaiyuan, et al.
Veröffentlicht: (2025)
Can LLMs Deeply Detect Complex Malicious Queries? A Framework for Jailbreaking via Obfuscating Intent
von: Shang, Shang, et al.
Veröffentlicht: (2024)
von: Shang, Shang, et al.
Veröffentlicht: (2024)
Output Supervision Can Obfuscate the Chain of Thought
von: Drori, Jacob, et al.
Veröffentlicht: (2025)
von: Drori, Jacob, et al.
Veröffentlicht: (2025)
EIM-TRNG: Obfuscating Deep Neural Network Weights with Encoding-in-Memory True Random Number Generator via RowHammer
von: Zhou, Ranyang, et al.
Veröffentlicht: (2025)
von: Zhou, Ranyang, et al.
Veröffentlicht: (2025)
An Empirical Study of Code Obfuscation Practices in the Google Play Store
von: Niroshan, Akila, et al.
Veröffentlicht: (2025)
von: Niroshan, Akila, et al.
Veröffentlicht: (2025)
TT-TFHE: a Torus Fully Homomorphic Encryption-Friendly Neural Network Architecture
von: Benamira, Adrien, et al.
Veröffentlicht: (2023)
von: Benamira, Adrien, et al.
Veröffentlicht: (2023)
Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reverse Engineering AI Agents
von: Crawford, Brian, et al.
Veröffentlicht: (2026)
von: Crawford, Brian, et al.
Veröffentlicht: (2026)
Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning Approach
von: Etter, Brian, et al.
Veröffentlicht: (2024)
von: Etter, Brian, et al.
Veröffentlicht: (2024)
Joint Universal Adversarial Perturbations with Interpretations
von: Ning, Liang-bo, et al.
Veröffentlicht: (2024)
von: Ning, Liang-bo, et al.
Veröffentlicht: (2024)
Deconstructing Obfuscation: A four-dimensional framework for evaluating Large Language Models assembly code deobfuscation capabilities
von: Tkachenko, Anton, et al.
Veröffentlicht: (2025)
von: Tkachenko, Anton, et al.
Veröffentlicht: (2025)
Multi-Granular Discretization for Interpretable Generalization in Precise Cyberattack Identification
von: Chung, Wen-Cheng, et al.
Veröffentlicht: (2025)
von: Chung, Wen-Cheng, et al.
Veröffentlicht: (2025)
Safeguarding AI Agents: Developing and Analyzing Safety Architectures
von: Domkundwar, Ishaan, et al.
Veröffentlicht: (2024)
von: Domkundwar, Ishaan, et al.
Veröffentlicht: (2024)
Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security
von: Chua, Gabriel
Veröffentlicht: (2025)
von: Chua, Gabriel
Veröffentlicht: (2025)
Enabling Regulatory Multi-Agent Collaboration: Architecture, Challenges, and Solutions
von: Hu, Qinnan, et al.
Veröffentlicht: (2025)
von: Hu, Qinnan, et al.
Veröffentlicht: (2025)
Backdoor Attack with Invisible Triggers Based on Model Architecture Modification
von: Ma, Yuan, et al.
Veröffentlicht: (2024)
von: Ma, Yuan, et al.
Veröffentlicht: (2024)
EXPLICATE: Enhancing Phishing Detection through Explainable AI and LLM-Powered Interpretability
von: Lim, Bryan, et al.
Veröffentlicht: (2025)
von: Lim, Bryan, et al.
Veröffentlicht: (2025)
Decoding BACnet Packets: A Large Language Model Approach for Packet Interpretation
von: Sharma, Rashi, et al.
Veröffentlicht: (2024)
von: Sharma, Rashi, et al.
Veröffentlicht: (2024)
An Interpretable Generalization Mechanism for Accurately Detecting Anomaly and Identifying Networking Intrusion Techniques
von: Pai, Hao-Ting, et al.
Veröffentlicht: (2024)
von: Pai, Hao-Ting, et al.
Veröffentlicht: (2024)
Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures
von: Schwarz, Dominik
Veröffentlicht: (2025)
von: Schwarz, Dominik
Veröffentlicht: (2025)
GUARD-CAN: Graph-Understanding and Recurrent Architecture for CAN Anomaly Detection
von: Kim, Hyeong Seon, et al.
Veröffentlicht: (2025)
von: Kim, Hyeong Seon, et al.
Veröffentlicht: (2025)
Advanced Prediction of Hypersonic Missile Trajectories with CNN-LSTM-GRU Architectures
von: Baradaran, Amir Hossein
Veröffentlicht: (2025)
von: Baradaran, Amir Hossein
Veröffentlicht: (2025)
AI-Governed Agent Architecture for Web-Trustworthy Tokenization of Alternative Assets
von: Borjigin, Ailiya, et al.
Veröffentlicht: (2025)
von: Borjigin, Ailiya, et al.
Veröffentlicht: (2025)
Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
von: Li, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2026)
Agentic AI for Cybersecurity: A Meta-Cognitive Architecture for Governable Autonomy
von: Kojukhov, Andrei, et al.
Veröffentlicht: (2026)
von: Kojukhov, Andrei, et al.
Veröffentlicht: (2026)
AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents
von: Zhang, Yixiang, et al.
Veröffentlicht: (2026)
von: Zhang, Yixiang, et al.
Veröffentlicht: (2026)
CAM-LDS: Cyber Attack Manifestations for Automatic Interpretation of System Logs and Security Alerts
von: Landauer, Max, et al.
Veröffentlicht: (2026)
von: Landauer, Max, et al.
Veröffentlicht: (2026)
Caging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare
von: Maiti, Saikat
Veröffentlicht: (2026)
von: Maiti, Saikat
Veröffentlicht: (2026)
SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment
von: Lin, Xixun, et al.
Veröffentlicht: (2026)
von: Lin, Xixun, et al.
Veröffentlicht: (2026)
MemTrust: A Zero-Trust Architecture for Unified AI Memory System
von: Zhou, Xing, et al.
Veröffentlicht: (2026)
von: Zhou, Xing, et al.
Veröffentlicht: (2026)
Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture
von: Xiang, Rong
Veröffentlicht: (2026)
von: Xiang, Rong
Veröffentlicht: (2026)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
von: Guan, Shaowei, et al.
Veröffentlicht: (2025)
von: Guan, Shaowei, et al.
Veröffentlicht: (2025)
Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token
von: Zychlinski, Shaked, et al.
Veröffentlicht: (2025)
von: Zychlinski, Shaked, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
An End-to-End Homomorphically Encrypted Neural Network
von: Florencio, Marcos, et al.
Veröffentlicht: (2025) -
Obfuscated Memory Malware Detection
von: P, Sharmila S, et al.
Veröffentlicht: (2024) -
Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
von: Wang, Ziwei, et al.
Veröffentlicht: (2026) -
Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice
von: Zhang, Gehao, et al.
Veröffentlicht: (2025) -
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
von: Zolkowski, Artur, et al.
Veröffentlicht: (2025)