Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schnabl, Christoph, Hugenroth, Daniel, Marino, Bill, Beresford, Alastair R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attestable Builds: Compiling Verifiable Binaries on Untrusted Systems using Trusted Execution Environments
von: Hugenroth, Daniel, et al.
Veröffentlicht: (2025)
von: Hugenroth, Daniel, et al.
Veröffentlicht: (2025)
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
von: Liu, Yule, et al.
Veröffentlicht: (2025)
von: Liu, Yule, et al.
Veröffentlicht: (2025)
Private, Verifiable, and Auditable AI Systems
von: South, Tobin
Veröffentlicht: (2025)
von: South, Tobin
Veröffentlicht: (2025)
SoK: Web Authentication in the Age of End-to-End Encryption
von: Blessing, Jenny, et al.
Veröffentlicht: (2024)
von: Blessing, Jenny, et al.
Veröffentlicht: (2024)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
von: Jin, Xisen, et al.
Veröffentlicht: (2026)
von: Jin, Xisen, et al.
Veröffentlicht: (2026)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
Giving AI Agents Access to Cryptocurrency and Smart Contracts Creates New Vectors of AI Harm
von: Marino, Bill, et al.
Veröffentlicht: (2025)
von: Marino, Bill, et al.
Veröffentlicht: (2025)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
von: Yang, Zhi, et al.
Veröffentlicht: (2026)
von: Yang, Zhi, et al.
Veröffentlicht: (2026)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
von: Das, Saswat, et al.
Veröffentlicht: (2025)
von: Das, Saswat, et al.
Veröffentlicht: (2025)
AuditGPT: Auditing Smart Contracts with ChatGPT
von: Xia, Shihao, et al.
Veröffentlicht: (2024)
von: Xia, Shihao, et al.
Veröffentlicht: (2024)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
von: Li, Tianhao, et al.
Veröffentlicht: (2024)
von: Li, Tianhao, et al.
Veröffentlicht: (2024)
SoK: Large Language Model Copyright Auditing via Fingerprinting
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
von: Shao, Shuo, et al.
Veröffentlicht: (2025)
Blueprints of Trust: AI System Cards for End to End Transparency and Governance
von: Sidhpurwala, Huzaifa, et al.
Veröffentlicht: (2025)
von: Sidhpurwala, Huzaifa, et al.
Veröffentlicht: (2025)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Dynamic Fog Computing for Enhanced LLM Execution in Medical Applications
von: Zagar, Philipp, et al.
Veröffentlicht: (2024)
von: Zagar, Philipp, et al.
Veröffentlicht: (2024)
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
von: Wang, Boxin, et al.
Veröffentlicht: (2023)
von: Wang, Boxin, et al.
Veröffentlicht: (2023)
Graph in the Vault: Protecting Edge GNN Inference with Trusted Execution Environment
von: Ding, Ruyi, et al.
Veröffentlicht: (2025)
von: Ding, Ruyi, et al.
Veröffentlicht: (2025)
What Matters For Safety Alignment?
von: Li, Xing, et al.
Veröffentlicht: (2026)
von: Li, Xing, et al.
Veröffentlicht: (2026)
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
von: Weiss, Roy, et al.
Veröffentlicht: (2024)
von: Weiss, Roy, et al.
Veröffentlicht: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
von: Łucki, Jakub, et al.
Veröffentlicht: (2024)
An Independent Safety Evaluation of Kimi K2.5
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2026)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2026)
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
von: Xu, Chejian, et al.
Veröffentlicht: (2025)
von: Xu, Chejian, et al.
Veröffentlicht: (2025)
Internal Safety Collapse in Frontier Large Language Models
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
von: Wu, Yutao, et al.
Veröffentlicht: (2026)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
von: Liu, Songyang, et al.
Veröffentlicht: (2025)
von: Liu, Songyang, et al.
Veröffentlicht: (2025)
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
von: Lee, JoonHo, et al.
Veröffentlicht: (2025)
Intent Laundering: AI Safety Datasets Are Not What They Seem
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
Trustworthy AI: Safety, Bias, and Privacy -- A Survey
von: Fang, Xingli, et al.
Veröffentlicht: (2025)
von: Fang, Xingli, et al.
Veröffentlicht: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
von: Ying, Zonghao, et al.
Veröffentlicht: (2025)
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
von: Lu, Guoxin, et al.
Veröffentlicht: (2026)
von: Lu, Guoxin, et al.
Veröffentlicht: (2026)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
von: Li, Lijun, et al.
Veröffentlicht: (2024)
von: Li, Lijun, et al.
Veröffentlicht: (2024)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
von: Xu, Zhao, et al.
Veröffentlicht: (2024)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
von: Wang, Zijun, et al.
Veröffentlicht: (2026)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
von: Kumar, Anurakt, et al.
Veröffentlicht: (2024)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
von: Chae, Kyubyung, et al.
Veröffentlicht: (2025)
Out of the Cage: How Stochastic Parrots Win in Cyber Security Environments
von: Rigaki, Maria, et al.
Veröffentlicht: (2023)
von: Rigaki, Maria, et al.
Veröffentlicht: (2023)
TrustChain: A Blockchain Framework for Auditing and Verifying Aggregators in Decentralized Federated Learning
von: Hallaji, Ehsan, et al.
Veröffentlicht: (2025)
von: Hallaji, Ehsan, et al.
Veröffentlicht: (2025)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
von: Wang, Kun, et al.
Veröffentlicht: (2026)
von: Wang, Kun, et al.
Veröffentlicht: (2026)
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
von: Zhang, Junbo, et al.
Veröffentlicht: (2026)
von: Zhang, Junbo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Attestable Builds: Compiling Verifiable Binaries on Untrusted Systems using Trusted Execution Environments
von: Hugenroth, Daniel, et al.
Veröffentlicht: (2025) -
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
von: Liu, Yule, et al.
Veröffentlicht: (2025) -
Private, Verifiable, and Auditable AI Systems
von: South, Tobin
Veröffentlicht: (2025) -
SoK: Web Authentication in the Age of End-to-End Encryption
von: Blessing, Jenny, et al.
Veröffentlicht: (2024) -
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
von: Jin, Xisen, et al.
Veröffentlicht: (2026)