Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands
Fuente:
arXiv
Saved in:
| Main Authors: | Seth, Pratinav, Sankarapu, Vinay Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Gap in XAI-Why Reliable Metrics Matter for Explainability and Compliance
by: Seth, Pratinav, et al.
Published: (2025)
by: Seth, Pratinav, et al.
Published: (2025)
Interpretability as Alignment: Making Internal Understanding a Design Principle
by: Sengupta, Aadit, et al.
Published: (2025)
by: Sengupta, Aadit, et al.
Published: (2025)
Orion-Bix: Bi-Axial Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
Distilling Tabular Foundation Models for Structured Health Data
by: Tanna, Aditya, et al.
Published: (2026)
by: Tanna, Aditya, et al.
Published: (2026)
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
by: Tanna, Aditya, et al.
Published: (2025)
by: Tanna, Aditya, et al.
Published: (2025)
Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees
by: Tanna, Aditya, et al.
Published: (2026)
by: Tanna, Aditya, et al.
Published: (2026)
DLBacktrace: A Model Agnostic Explainability for any Deep Learning Models
by: Sankarapu, Vinay Kumar, et al.
Published: (2024)
by: Sankarapu, Vinay Kumar, et al.
Published: (2024)
xai_evals : A Framework for Evaluating Post-Hoc Local Explanation Methods
by: Seth, Pratinav, et al.
Published: (2025)
by: Seth, Pratinav, et al.
Published: (2025)
Ensembling Tabular Foundation Models - A Diversity Ceiling And A Calibration Trap
by: Tanna, Aditya, et al.
Published: (2026)
by: Tanna, Aditya, et al.
Published: (2026)
Data Presentation Over Architecture: Resampling Strategies for Credit Risk Prediction with Tabular Foundation Models
by: Tanna, Aditya, et al.
Published: (2026)
by: Tanna, Aditya, et al.
Published: (2026)
Interpretability-Aware Pruning for Efficient Medical Image Analysis
by: Malik, Nikita, et al.
Published: (2025)
by: Malik, Nikita, et al.
Published: (2025)
Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution
by: Sadhu, Saisab, et al.
Published: (2026)
by: Sadhu, Saisab, et al.
Published: (2026)
Exploring Fine-Tuning for Tabular Foundation Models
by: Tanna, Aditya, et al.
Published: (2026)
by: Tanna, Aditya, et al.
Published: (2026)
Beyond KL Divergence: Policy Optimization with Flexible Bregman Divergences for LLM Reasoning
by: Yuan, Rui, et al.
Published: (2026)
by: Yuan, Rui, et al.
Published: (2026)
Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
by: Khandoga, Mykola, et al.
Published: (2026)
by: Khandoga, Mykola, et al.
Published: (2026)
Shaping the Prior: How Synthetic Task Distributions Determine Tabular Foundation Model Quality
by: Bouadi, Mohamed, et al.
Published: (2026)
by: Bouadi, Mohamed, et al.
Published: (2026)
$C$-$ΔΘ$: Circuit-Restricted Weight Arithmetic for Selective Refusal
by: Kasliwal, Aditya, et al.
Published: (2026)
by: Kasliwal, Aditya, et al.
Published: (2026)
Position: State-of-the-Art Claims Require State-of-the-Art Evidence
by: Oh, YongKyung
Published: (2026)
by: Oh, YongKyung
Published: (2026)
Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
by: Cho, Dongkyu Derek, et al.
Published: (2025)
by: Cho, Dongkyu Derek, et al.
Published: (2025)
Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
by: Kumar, Anantha Padmanaban Krishna
Published: (2025)
Towards Developing Safety Assurance Cases for Learning-Enabled Medical Cyber-Physical Systems
by: Bagheri, Maryam, et al.
Published: (2022)
by: Bagheri, Maryam, et al.
Published: (2022)
Relational GNNs Cannot Learn $C_2$ Features for Planning
by: Chen, Dillon Z.
Published: (2025)
by: Chen, Dillon Z.
Published: (2025)
ART: Adaptive Reasoning Trees for Explainable Claim Verification
by: Wadhwa, Sahil, et al.
Published: (2026)
by: Wadhwa, Sahil, et al.
Published: (2026)
Verified Neural Compressed Sensing
by: Bunel, Rudy, et al.
Published: (2024)
by: Bunel, Rudy, et al.
Published: (2024)
Co-Activation Graph Analysis of Safety-Verified and Explainable Deep Reinforcement Learning Policies
by: Gross, Dennis, et al.
Published: (2025)
by: Gross, Dennis, et al.
Published: (2025)
xAI-Drop: Don't Use What You Cannot Explain
by: De Luca, Vincenzo Marco, et al.
Published: (2024)
by: De Luca, Vincenzo Marco, et al.
Published: (2024)
From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems
by: Stefani, Thomas, et al.
Published: (2026)
by: Stefani, Thomas, et al.
Published: (2026)
AlignTune: Modular Toolkit for Post-Training Alignment of Large Language Models
by: Lyngkhoi, R E Zera Marveen, et al.
Published: (2026)
by: Lyngkhoi, R E Zera Marveen, et al.
Published: (2026)
Text-to-Image Diffusion Models Cannot Count, and Prompt Refinement Cannot Help
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
Operationalising Artificial Intelligence Bills of Materials (AIBOMs) for Verifiable AI Provenance and Lifecycle Assurance
by: Radanliev, Petar, et al.
Published: (2026)
by: Radanliev, Petar, et al.
Published: (2026)
OpenKedge: Governing Agentic Mutation with Execution-Bound Safety and Evidence Chains
by: He, Jun, et al.
Published: (2026)
by: He, Jun, et al.
Published: (2026)
Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
Edge-FIT: Federated Instruction Tuning of Quantized LLMs for Privacy-Preserving Smart Home Environments
by: Venkatesh, Vinay, et al.
Published: (2025)
by: Venkatesh, Vinay, et al.
Published: (2025)
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
by: Cai, Xin-Qiang, et al.
Published: (2025)
by: Cai, Xin-Qiang, et al.
Published: (2025)
Behaviour Distillation
by: Lupu, Andrei, et al.
Published: (2024)
by: Lupu, Andrei, et al.
Published: (2024)
Curriculum Learning for Safety Alignment
by: Kumar, Sandeep, et al.
Published: (2026)
by: Kumar, Sandeep, et al.
Published: (2026)
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
Similar Items
-
Bridging the Gap in XAI-Why Reliable Metrics Matter for Explainability and Compliance
by: Seth, Pratinav, et al.
Published: (2025) -
Interpretability as Alignment: Making Internal Understanding a Design Principle
by: Sengupta, Aadit, et al.
Published: (2025) -
Orion-Bix: Bi-Axial Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025) -
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025) -
Distilling Tabular Foundation Models for Structured Health Data
by: Tanna, Aditya, et al.
Published: (2026)