Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
Fuente:
arXiv
Guardado en:
| Autor principal: | Scrivens, Arsenios |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
por: Scrivens, Arsenios
Publicado: (2026)
por: Scrivens, Arsenios
Publicado: (2026)
Holographic Invariant Storage: Design-Time Safety Contracts via Vector Symbolic Architectures
por: Scrivens, Arsenios
Publicado: (2026)
por: Scrivens, Arsenios
Publicado: (2026)
NeuroAI for AI Safety
por: Mineault, Patrick, et al.
Publicado: (2024)
por: Mineault, Patrick, et al.
Publicado: (2024)
Bridging Efficiency and Safety: Formal Verification of Neural Networks with Early Exits
por: Elboher, Yizhak Yisrael, et al.
Publicado: (2025)
por: Elboher, Yizhak Yisrael, et al.
Publicado: (2025)
TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift
por: Mahato, Dipesh Tharu, et al.
Publicado: (2025)
por: Mahato, Dipesh Tharu, et al.
Publicado: (2025)
Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce Grokking
por: Lyu, Kaifeng, et al.
Publicado: (2023)
por: Lyu, Kaifeng, et al.
Publicado: (2023)
Divergence of Empirical Neural Tangent Kernel in Classification Problems
por: Yu, Zixiong, et al.
Publicado: (2025)
por: Yu, Zixiong, et al.
Publicado: (2025)
Bridging Neural ODE and ResNet: A Formal Error Bound for Safety Verification
por: Sayed, Abdelrahman Sayed, et al.
Publicado: (2025)
por: Sayed, Abdelrahman Sayed, et al.
Publicado: (2025)
Shutdown Safety Valves for Advanced AI
por: Conitzer, Vincent
Publicado: (2026)
por: Conitzer, Vincent
Publicado: (2026)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
por: Griffin, Charlie, et al.
Publicado: (2024)
por: Griffin, Charlie, et al.
Publicado: (2024)
Holistic Safety and Responsibility Evaluations of Advanced AI Models
por: Weidinger, Laura, et al.
Publicado: (2024)
por: Weidinger, Laura, et al.
Publicado: (2024)
Self-Improving Safety Performance of Reinforcement Learning Based Driving with Black-Box Verification Algorithms
por: Dagdanov, Resul, et al.
Publicado: (2022)
por: Dagdanov, Resul, et al.
Publicado: (2022)
Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability
por: Liu, Yushen, et al.
Publicado: (2026)
por: Liu, Yushen, et al.
Publicado: (2026)
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
por: Liu, Chenruo, et al.
Publicado: (2025)
por: Liu, Chenruo, et al.
Publicado: (2025)
Computational Safety for Generative AI: A Signal Processing Perspective
por: Chen, Pin-Yu
Publicado: (2025)
por: Chen, Pin-Yu
Publicado: (2025)
Effective Controllable Bias Mitigation for Classification and Retrieval using Gate Adapters
por: Masoudian, Shahed, et al.
Publicado: (2024)
por: Masoudian, Shahed, et al.
Publicado: (2024)
Spectral Gating Networks
por: Zhang, Jusheng, et al.
Publicado: (2026)
por: Zhang, Jusheng, et al.
Publicado: (2026)
Does Faithfulness Conflict with Plausibility? An Empirical Study in Explainable AI across NLP Tasks
por: Lu, Xiaolei, et al.
Publicado: (2024)
por: Lu, Xiaolei, et al.
Publicado: (2024)
Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
por: Martra, Pere
Publicado: (2025)
por: Martra, Pere
Publicado: (2025)
NASP-T: A Fuzzy Neuro-Symbolic Transformer for Logic-Constrained Aviation Safety Report Classification
por: Machot, Fadi Al, et al.
Publicado: (2025)
por: Machot, Fadi Al, et al.
Publicado: (2025)
A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety
por: Lee, Hyunin, et al.
Publicado: (2024)
por: Lee, Hyunin, et al.
Publicado: (2024)
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
por: Overman, William, et al.
Publicado: (2025)
por: Overman, William, et al.
Publicado: (2025)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
por: Korbak, Tomek, et al.
Publicado: (2025)
por: Korbak, Tomek, et al.
Publicado: (2025)
DB-FGA-Net: Dual Backbone Frequency Gated Attention Network for Multi-Class Brain Tumor Classification with Grad-CAM Interpretability
por: Shreya, Saraf Anzum, et al.
Publicado: (2025)
por: Shreya, Saraf Anzum, et al.
Publicado: (2025)
International AI Safety Report
por: Bengio, Yoshua, et al.
Publicado: (2025)
por: Bengio, Yoshua, et al.
Publicado: (2025)
Domain Gating Ensemble Networks for AI-Generated Text Detection
por: Tripathi, Arihant, et al.
Publicado: (2025)
por: Tripathi, Arihant, et al.
Publicado: (2025)
xEEGNet: Towards Explainable AI in EEG Dementia Classification
por: Zanola, Andrea, et al.
Publicado: (2025)
por: Zanola, Andrea, et al.
Publicado: (2025)
Lightweight Safety Classification Using Pruned Language Models
por: Sawtell, Mason, et al.
Publicado: (2024)
por: Sawtell, Mason, et al.
Publicado: (2024)
Green MLOps to Green GenOps: An Empirical Study of Energy Consumption in Discriminative and Generative AI Operations
por: Sánchez-Mompó, Adrián, et al.
Publicado: (2025)
por: Sánchez-Mompó, Adrián, et al.
Publicado: (2025)
Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided Gate
por: Zheng, Liangwei Nathan, et al.
Publicado: (2025)
por: Zheng, Liangwei Nathan, et al.
Publicado: (2025)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
por: Li, Guanlin, et al.
Publicado: (2025)
por: Li, Guanlin, et al.
Publicado: (2025)
Introduction to AI Safety, Ethics, and Society
por: Hendrycks, Dan
Publicado: (2024)
por: Hendrycks, Dan
Publicado: (2024)
Spatially-Delineated Domain-Adapted AI Classification: An Application for Oncology Data
por: Farhadloo, Majid, et al.
Publicado: (2025)
por: Farhadloo, Majid, et al.
Publicado: (2025)
DynamicGate MLP Conditional Computation via Learned Structural Dropout and Input Dependent Gating for Functional Plasticity
por: Choi, Yong Il
Publicado: (2026)
por: Choi, Yong Il
Publicado: (2026)
Generative AI Agents in Autonomous Machines: A Safety Perspective
por: Jabbour, Jason, et al.
Publicado: (2024)
por: Jabbour, Jason, et al.
Publicado: (2024)
Automated Conjecture Resolution with Formal Verification
por: Ju, Haocheng, et al.
Publicado: (2026)
por: Ju, Haocheng, et al.
Publicado: (2026)
Neural Network Verification with PyRAT
por: Lemesle, Augustin, et al.
Publicado: (2024)
por: Lemesle, Augustin, et al.
Publicado: (2024)
From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems
por: Stefani, Thomas, et al.
Publicado: (2026)
por: Stefani, Thomas, et al.
Publicado: (2026)
Fitting Multilinear Polynomials for Logic Gate Networks
por: Kim, Youngsung
Publicado: (2026)
por: Kim, Youngsung
Publicado: (2026)
Uncertainty Estimation using Variance-Gated Distributions
por: Gillis, H. Martin, et al.
Publicado: (2025)
por: Gillis, H. Martin, et al.
Publicado: (2025)
Ejemplares similares
-
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
por: Scrivens, Arsenios
Publicado: (2026) -
Holographic Invariant Storage: Design-Time Safety Contracts via Vector Symbolic Architectures
por: Scrivens, Arsenios
Publicado: (2026) -
NeuroAI for AI Safety
por: Mineault, Patrick, et al.
Publicado: (2024) -
Bridging Efficiency and Safety: Formal Verification of Neural Networks with Early Exits
por: Elboher, Yizhak Yisrael, et al.
Publicado: (2025) -
TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift
por: Mahato, Dipesh Tharu, et al.
Publicado: (2025)