TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift
Fuente:
arXiv
Saved in:
| Main Authors: | Mahato, Dipesh Tharu, Poudel, Rohan, Dhungana, Pramod |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep learning-based fault identification in condition monitoring
by: Dhungana, Hariom, et al.
Published: (2024)
by: Dhungana, Hariom, et al.
Published: (2024)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety
by: Suleymanov, Umid, et al.
Published: (2026)
by: Suleymanov, Umid, et al.
Published: (2026)
SafeSeek: Universal Attribution of Safety Circuits in Language Models
by: Yu, Miao, et al.
Published: (2026)
by: Yu, Miao, et al.
Published: (2026)
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
by: Wang, Tongxi, et al.
Published: (2026)
by: Wang, Tongxi, et al.
Published: (2026)
DriftXpress: Faster Drifting Models via Projected RKHS Fields
by: Falahati, Ali, et al.
Published: (2026)
by: Falahati, Ali, et al.
Published: (2026)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
by: Oh, Sejoon, et al.
Published: (2024)
by: Oh, Sejoon, et al.
Published: (2024)
Empirical Validation of the Classification-Verification Dichotomy for AI Safety Gates
by: Scrivens, Arsenios
Published: (2026)
by: Scrivens, Arsenios
Published: (2026)
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
by: Scrivens, Arsenios
Published: (2026)
by: Scrivens, Arsenios
Published: (2026)
Guarding the Meaning: Self-Supervised Training for Semantic Robustness in Guard Models
by: Pinneri, Cristina, et al.
Published: (2025)
by: Pinneri, Cristina, et al.
Published: (2025)
Bridging Efficiency and Safety: Formal Verification of Neural Networks with Early Exits
by: Elboher, Yizhak Yisrael, et al.
Published: (2025)
by: Elboher, Yizhak Yisrael, et al.
Published: (2025)
Stateful KV Cache Management for LLMs: Balancing Space, Time, Accuracy, and Positional Fidelity
by: Poudel, Pratik
Published: (2025)
by: Poudel, Pratik
Published: (2025)
Black-box Adversarial Attacks on Network-wide Multi-step Traffic State Prediction Models
by: Poudel, Bibek, et al.
Published: (2021)
by: Poudel, Bibek, et al.
Published: (2021)
TriP-LLM: A Tri-Branch Patch-wise Large Language Model Framework for Time-Series Anomaly Detection
by: Yu, Yuan-Cheng, et al.
Published: (2025)
by: Yu, Yuan-Cheng, et al.
Published: (2025)
On the Wasserstein Gradient Flow Interpretation of Drifting Models
by: Gretton, Arthur, et al.
Published: (2026)
by: Gretton, Arthur, et al.
Published: (2026)
Bridging Neural ODE and ResNet: A Formal Error Bound for Safety Verification
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
by: Sayed, Abdelrahman Sayed, et al.
Published: (2025)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
TARo: Token-level Adaptive Routing for LLM Test-time Alignment
by: Rai, Arushi, et al.
Published: (2026)
by: Rai, Arushi, et al.
Published: (2026)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
ETAGE: Enhanced Test Time Adaptation with Integrated Entropy and Gradient Norms for Robust Model Performance
by: Shamsi, Afshar, et al.
Published: (2024)
by: Shamsi, Afshar, et al.
Published: (2024)
Test-Time Training Undermines Safety Guardrails
by: Antonelli, Simone, et al.
Published: (2026)
by: Antonelli, Simone, et al.
Published: (2026)
Tri-Level Navigator: LLM-Empowered Tri-Level Learning for Time Series OOD Generalization
by: Jian, Chengtao, et al.
Published: (2024)
by: Jian, Chengtao, et al.
Published: (2024)
Using Quality Attribute Scenarios for ML Model Test Case Generation
by: Brower-Sinning, Rachel, et al.
Published: (2024)
by: Brower-Sinning, Rachel, et al.
Published: (2024)
Drift Flow Matching
by: Ma, Chenrui, et al.
Published: (2026)
by: Ma, Chenrui, et al.
Published: (2026)
Drift Q-Learning
by: Houssaini, Anas, et al.
Published: (2026)
by: Houssaini, Anas, et al.
Published: (2026)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
by: Shao, Shuo, et al.
Published: (2024)
by: Shao, Shuo, et al.
Published: (2024)
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
by: Zhao, Chu, et al.
Published: (2026)
by: Zhao, Chu, et al.
Published: (2026)
An Innovative Next Activity Prediction Using Process Entropy and Dynamic Attribute-Wise-Transformer in Predictive Business Process Monitoring
by: Zare, Hadi, et al.
Published: (2025)
by: Zare, Hadi, et al.
Published: (2025)
Technical Report: Evaluating Goal Drift in Language Model Agents
by: Arike, Rauno, et al.
Published: (2025)
by: Arike, Rauno, et al.
Published: (2025)
SymDrift: One-Shot Generative Modeling under Symmetries
by: Darouich, Samir, et al.
Published: (2026)
by: Darouich, Samir, et al.
Published: (2026)
TLoRA: Tri-Matrix Low-Rank Adaptation of Large Language Models
by: Islam, Tanvir
Published: (2025)
by: Islam, Tanvir
Published: (2025)
Zero-Shot Attribution for Large Language Models: A Distribution Testing Approach
by: Canonne, Clément L., et al.
Published: (2025)
by: Canonne, Clément L., et al.
Published: (2025)
G-Drift MIA: Membership Inference via Gradient-Induced Feature Drift in LLMs
by: Ranjan, Ravi, et al.
Published: (2026)
by: Ranjan, Ravi, et al.
Published: (2026)
FlakyGuard: Automatically Fixing Flaky Tests at Industry Scale
by: Li, Chengpeng, et al.
Published: (2025)
by: Li, Chengpeng, et al.
Published: (2025)
CALM: A Causal Analysis Language Model for Tabular Data in Complex Systems with Local Scores, Conditional Independence Tests, and Relation Attributes
by: Fan, Zhenjiang, et al.
Published: (2025)
by: Fan, Zhenjiang, et al.
Published: (2025)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
by: Zhan, Weixiao, et al.
Published: (2026)
by: Zhan, Weixiao, et al.
Published: (2026)
Guarding Graph Neural Networks for Unsupervised Graph Anomaly Detection
by: Bei, Yuanchen, et al.
Published: (2024)
by: Bei, Yuanchen, et al.
Published: (2024)
Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI
by: Advani, Laksh
Published: (2026)
by: Advani, Laksh
Published: (2026)
Similar Items
-
Deep learning-based fault identification in condition monitoring
by: Dhungana, Hariom, et al.
Published: (2024) -
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026) -
CourtGuard: A Model-Agnostic Framework for Zero-Shot Policy Adaptation in LLM Safety
by: Suleymanov, Umid, et al.
Published: (2026) -
SafeSeek: Universal Attribution of Safety Circuits in Language Models
by: Yu, Miao, et al.
Published: (2026) -
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
by: Hossain, Ismail, et al.
Published: (2026)