Trustless Audits without Revealing Data or Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Waiwitlikhit, Suppakit, Stoica, Ion, Sun, Yi, Hashimoto, Tatsunori, Kang, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model
por: Bicakci, Kemal
Publicado: (2026)
por: Bicakci, Kemal
Publicado: (2026)
Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric
por: Shvartzshnaider, Yan, et al.
Publicado: (2024)
por: Shvartzshnaider, Yan, et al.
Publicado: (2024)
MPC-Minimized Secure LLM Inference
por: Rathee, Deevashwer, et al.
Publicado: (2024)
por: Rathee, Deevashwer, et al.
Publicado: (2024)
Differentially Private Data Release on Graphs: Inefficiencies and Unfairness
por: Fioretto, Ferdinando, et al.
Publicado: (2024)
por: Fioretto, Ferdinando, et al.
Publicado: (2024)
Machine Unlearning Fails to Remove Data Poisoning Attacks
por: Pawelczyk, Martin, et al.
Publicado: (2024)
por: Pawelczyk, Martin, et al.
Publicado: (2024)
Auditing Prompt Caching in Language Model APIs
por: Gu, Chenchen, et al.
Publicado: (2025)
por: Gu, Chenchen, et al.
Publicado: (2025)
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
por: Jang, Yeonwoo, et al.
Publicado: (2025)
por: Jang, Yeonwoo, et al.
Publicado: (2025)
Unifying Re-Identification, Attribute Inference, and Data Reconstruction Risks in Differential Privacy
por: Kulynych, Bogdan, et al.
Publicado: (2025)
por: Kulynych, Bogdan, et al.
Publicado: (2025)
Secure Multi-Modal Data Fusion in Federated Digital Health Systems via MCP
por: Aueawatthanaphisut, Aueaphum
Publicado: (2025)
por: Aueawatthanaphisut, Aueaphum
Publicado: (2025)
JSTprove: Pioneering Verifiable AI for a Trustless Future
por: Gold, Jonathan, et al.
Publicado: (2025)
por: Gold, Jonathan, et al.
Publicado: (2025)
Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
por: Rabanser, Stephan, et al.
Publicado: (2025)
por: Rabanser, Stephan, et al.
Publicado: (2025)
Robust Safety Monitoring of Language Models via Activation Watermarking
por: Aremu, Toluwani, et al.
Publicado: (2026)
por: Aremu, Toluwani, et al.
Publicado: (2026)
A Survey of Privacy-Preserving Model Explanations: Privacy Risks, Attacks, and Countermeasures
por: Nguyen, Thanh Tam, et al.
Publicado: (2024)
por: Nguyen, Thanh Tam, et al.
Publicado: (2024)
Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations
por: Ezzeddine, Fatima, et al.
Publicado: (2024)
por: Ezzeddine, Fatima, et al.
Publicado: (2024)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
por: Barrett, Anthony M., et al.
Publicado: (2024)
por: Barrett, Anthony M., et al.
Publicado: (2024)
Auditing Pay-Per-Token in Large Language Models
por: Velasco, Ander Artola, et al.
Publicado: (2025)
por: Velasco, Ander Artola, et al.
Publicado: (2025)
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
por: Tan, Zhen, et al.
Publicado: (2024)
por: Tan, Zhen, et al.
Publicado: (2024)
Inferring Discussion Topics about Exploitation of Vulnerabilities from Underground Hacking Forums
por: Moreno-Vera, Felipe
Publicado: (2024)
por: Moreno-Vera, Felipe
Publicado: (2024)
Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers
por: Burgess, Mark A., et al.
Publicado: (2024)
por: Burgess, Mark A., et al.
Publicado: (2024)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
por: Shao, Minghao, et al.
Publicado: (2024)
por: Shao, Minghao, et al.
Publicado: (2024)
SoK: On the Offensive Potential of AI
por: Schröer, Saskia Laura, et al.
Publicado: (2024)
por: Schröer, Saskia Laura, et al.
Publicado: (2024)
SecGenAI: Enhancing Security of Cloud-based Generative AI Applications within Australian Critical Technologies of National Interest
por: Haryanto, Christoforus Yoga, et al.
Publicado: (2024)
por: Haryanto, Christoforus Yoga, et al.
Publicado: (2024)
Privacy at a Price: Exploring its Dual Impact on AI Fairness
por: Yang, Mengmeng, et al.
Publicado: (2024)
por: Yang, Mengmeng, et al.
Publicado: (2024)
PUFFLE: Balancing Privacy, Utility, and Fairness in Federated Learning
por: Corbucci, Luca, et al.
Publicado: (2024)
por: Corbucci, Luca, et al.
Publicado: (2024)
Adversarial Augmentation and Active Sampling for Robust Cyber Anomaly Detection
por: Benabderrahmane, Sidahmed, et al.
Publicado: (2025)
por: Benabderrahmane, Sidahmed, et al.
Publicado: (2025)
Password-Activated Shutdown Protocols for Misaligned Frontier Agents
por: Williams, Kai, et al.
Publicado: (2025)
por: Williams, Kai, et al.
Publicado: (2025)
Decentralized autonomous organization and blockchain-based incentivization framework for community-based facilities management
por: Ly, Reachsak, et al.
Publicado: (2026)
por: Ly, Reachsak, et al.
Publicado: (2026)
Watermarking Should Be Treated as a Monitoring Primitive
por: Aremu, Toluwani, et al.
Publicado: (2026)
por: Aremu, Toluwani, et al.
Publicado: (2026)
The New Frontier of Cybersecurity: Emerging Threats and Innovations
por: Dave, Daksh, et al.
Publicado: (2023)
por: Dave, Daksh, et al.
Publicado: (2023)
Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks
por: Di, Jimmy Z., et al.
Publicado: (2022)
por: Di, Jimmy Z., et al.
Publicado: (2022)
Robustness and Cybersecurity in the EU Artificial Intelligence Act
por: Nolte, Henrik, et al.
Publicado: (2025)
por: Nolte, Henrik, et al.
Publicado: (2025)
A Public Theory of Distillation Resistance via Constraint-Coupled Reasoning Architectures
por: Wei, Peng, et al.
Publicado: (2026)
por: Wei, Peng, et al.
Publicado: (2026)
FAIRPLAI: A Human-in-the-Loop Approach to Fair and Private Machine Learning
por: Sanchez Jr., David, et al.
Publicado: (2025)
por: Sanchez Jr., David, et al.
Publicado: (2025)
VISION: Robust and Interpretable Code Vulnerability Detection Leveraging Counterfactual Augmentation
por: Egea, David, et al.
Publicado: (2025)
por: Egea, David, et al.
Publicado: (2025)
An In-Depth Investigation of Data Collection in LLM App Ecosystems
por: Wu, Yuhao, et al.
Publicado: (2024)
por: Wu, Yuhao, et al.
Publicado: (2024)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
por: Zhang, Andy K., et al.
Publicado: (2024)
por: Zhang, Andy K., et al.
Publicado: (2024)
Black-Box Access is Insufficient for Rigorous AI Audits
por: Casper, Stephen, et al.
Publicado: (2024)
por: Casper, Stephen, et al.
Publicado: (2024)
Private, Verifiable, and Auditable AI Systems
por: South, Tobin
Publicado: (2025)
por: South, Tobin
Publicado: (2025)
Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
por: Wu, Yixin, et al.
Publicado: (2025)
por: Wu, Yixin, et al.
Publicado: (2025)
Privacy Auditing of Large Language Models
por: Panda, Ashwinee, et al.
Publicado: (2025)
por: Panda, Ashwinee, et al.
Publicado: (2025)
Ejemplares similares
-
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model
por: Bicakci, Kemal
Publicado: (2026) -
Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric
por: Shvartzshnaider, Yan, et al.
Publicado: (2024) -
MPC-Minimized Secure LLM Inference
por: Rathee, Deevashwer, et al.
Publicado: (2024) -
Differentially Private Data Release on Graphs: Inefficiencies and Unfairness
por: Fioretto, Ferdinando, et al.
Publicado: (2024) -
Machine Unlearning Fails to Remove Data Poisoning Attacks
por: Pawelczyk, Martin, et al.
Publicado: (2024)