Skip to content
Descubridor Institucional UMAR
Inicio
Búsqueda avanzada
Explorar
Inicio
Búsqueda avanzada
Explorar
Login
Language
English
Deutsch
Español
Français
Italiano
All Fields
Title
Author
Subject
Call Number
ISBN/ISSN
Tag
Find
Advanced
Unmasking the Trojan Horse: Provable Detection and Mitigation of Deceptive Alignment in AI Systems
Unmasking the Trojan Horse: Provable Detection and Mitigation of Deceptive Alignment in AI Systems
Fuente:
Zenodo
Saved in:
Bibliographic Details
Main Authors:
Revista, Zen
,
IA, 10
Format:
Recurso digital
Published:
Zenodo
2025
Online Access:
Acceder al recurso
Tags:
Add Tag
No Tags, Be the first to tag this record!
Cite this
Text this
Email this
Print
Export Record
Export to RefWorks
Export to EndNoteWeb
Export to EndNote
Save to List
Permanent link
Holdings
Description
Comments
Similar Items
Staff View
Internet
https://doi.org/10.5281/zenodo.17795765
Similar Items
Emergent Goal Concordance: Architecting Self-Correcting AI Alignment Systems
by: Revista, Zen, et al.
Published: (2025)
Unmasking Causal Representations: A Novel Masking Paradigm for Explainable Language Models
by: Revista, Zen, et al.
Published: (2025)
Data Quality Cascades in AI: Characterizing, Mitigating, and Preventing Error Propagation
by: Revista, Zen, et al.
Published: (2025)
The AI Alignment Loop: Iterative Self-Correction for Robust and Generalizable Models
by: Revista, Zen, et al.
Published: (2025)
Normative Engineering of AI: Leveraging Inner Product Geometry for Robustness and Alignment
by: Revista, Zen, et al.
Published: (2025)