Sensitivity Uncertainty Alignment in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hiremath, Prakul Sunil, Hiremath, Harshit R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Hiremath Early Detection (HED) Score: A Measure-Theoretic Evaluation Standard for Temporal Intelligence
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
A High-Recall Cost-Sensitive Machine Learning Framework for Real-Time Online Banking Transaction Fraud Detection
von: R., Karthikeyan V., et al.
Veröffentlicht: (2026)
von: R., Karthikeyan V., et al.
Veröffentlicht: (2026)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
von: Chakraborty, Amit, et al.
Veröffentlicht: (2025)
von: Chakraborty, Amit, et al.
Veröffentlicht: (2025)
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
von: Yu, Hao, et al.
Veröffentlicht: (2026)
von: Yu, Hao, et al.
Veröffentlicht: (2026)
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
von: Hoang, Tien Dat
Veröffentlicht: (2025)
von: Hoang, Tien Dat
Veröffentlicht: (2025)
SRFed: Mitigating Poisoning Attacks in Privacy-Preserving Federated Learning with Heterogeneous Data
von: Lu, Yiwen
Veröffentlicht: (2026)
von: Lu, Yiwen
Veröffentlicht: (2026)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
von: Nathanson, Samuel, et al.
Veröffentlicht: (2025)
von: Nathanson, Samuel, et al.
Veröffentlicht: (2025)
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
von: Collu, Matteo Gioele, et al.
Veröffentlicht: (2023)
von: Collu, Matteo Gioele, et al.
Veröffentlicht: (2023)
Automated Hardware Trojan Insertion in Industrial-Scale Designs
von: Popryho, Yaroslav, et al.
Veröffentlicht: (2025)
von: Popryho, Yaroslav, et al.
Veröffentlicht: (2025)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
von: Young, Richard J., et al.
Veröffentlicht: (2026)
von: Young, Richard J., et al.
Veröffentlicht: (2026)
Photonic AI: A Hybrid Diffractive Holographic Neural System for Passive Optical Real-Time Image Classification
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
Network Structures as an Attack Surface: Topology-Based Privacy Leakage in Federated Learning
von: Rangwala, Murtaza, et al.
Veröffentlicht: (2025)
von: Rangwala, Murtaza, et al.
Veröffentlicht: (2025)
Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning
von: Subedar, Noah, et al.
Veröffentlicht: (2025)
von: Subedar, Noah, et al.
Veröffentlicht: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
von: Azov, Guy, et al.
Veröffentlicht: (2026)
von: Azov, Guy, et al.
Veröffentlicht: (2026)
Density-aware Sample-specific Attack
von: Wang, Qiyuan, et al.
Veröffentlicht: (2026)
von: Wang, Qiyuan, et al.
Veröffentlicht: (2026)
Towards Low-Latency and Adaptive Ransomware Detection Using Contrastive Learning
von: Pan, Zhixin, et al.
Veröffentlicht: (2025)
von: Pan, Zhixin, et al.
Veröffentlicht: (2025)
Generalizable and Interpretable RF Fingerprinting with Shapelet-Enhanced Large Language Models
von: Zhao, Tianya, et al.
Veröffentlicht: (2026)
von: Zhao, Tianya, et al.
Veröffentlicht: (2026)
GIRL: Generative Imagination Reinforcement Learning via Information-Theoretic Hallucination Control
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
Regret-Aware Policy Optimization: Environment-Level Memory for Replay Suppression under Delayed Harm
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
PACT: Reducing Alert Fatigue in Low-Prevalence SOC Streams with Triggered Active Learning
von: Ndichu, Samuel, et al.
Veröffentlicht: (2026)
von: Ndichu, Samuel, et al.
Veröffentlicht: (2026)
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity
von: Rashidi, Mohammadreza
Veröffentlicht: (2026)
von: Rashidi, Mohammadreza
Veröffentlicht: (2026)
Risk-Calibrated Bayesian Streaming Intrusion Detection with SRE-Aligned Decisions
von: Youssef, Michel
Veröffentlicht: (2025)
von: Youssef, Michel
Veröffentlicht: (2025)
A Protocol-Language Model for Network Intrusion (Without Deep Packet Inspection)
von: Sharma, Vivek Kumar
Veröffentlicht: (2026)
von: Sharma, Vivek Kumar
Veröffentlicht: (2026)
Security Considerations for Multi-agent Systems
von: Nguyen, Tam, et al.
Veröffentlicht: (2026)
von: Nguyen, Tam, et al.
Veröffentlicht: (2026)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
von: Dawson, Ads, et al.
Veröffentlicht: (2025)
von: Dawson, Ads, et al.
Veröffentlicht: (2025)
PARD-SSM: Probabilistic Cyber-Attack Regime Detection via Variational Switching State-Space Models
von: Hiremath, Prakul Sunil, et al.
Veröffentlicht: (2026)
von: Hiremath, Prakul Sunil, et al.
Veröffentlicht: (2026)
Learning Nonlinear Regime Transitions via Semi-Parametric State-Space Models
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026)
Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
Breaking Boundaries: Balancing Performance and Robustness in Deep Wireless Traffic Forecasting
von: Ilbert, Romain, et al.
Veröffentlicht: (2023)
von: Ilbert, Romain, et al.
Veröffentlicht: (2023)
SCAFDS: Edge-Feature Graph Attention for Interbank Fraud Detection with Attribution-Grounded SAR Generation
von: Uddin, Mohammad Nasir
Veröffentlicht: (2026)
von: Uddin, Mohammad Nasir
Veröffentlicht: (2026)
Can Graph-Based Microservice Performance Detection Be Used for Microservice Intrusion Detection?
von: Ma, Yunjian
Veröffentlicht: (2026)
von: Ma, Yunjian
Veröffentlicht: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
Static Attribution of Android Residential Proxy Malware Using Graph Kernels
von: Clark, Peter, et al.
Veröffentlicht: (2026)
von: Clark, Peter, et al.
Veröffentlicht: (2026)
Non-Adaptive Adversarial Face Generation
von: Kim, Sunpill, et al.
Veröffentlicht: (2025)
von: Kim, Sunpill, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption
von: Morales, Jaime, et al.
Veröffentlicht: (2026)
von: Morales, Jaime, et al.
Veröffentlicht: (2026)
Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning
von: Huff, Philip, et al.
Veröffentlicht: (2026)
von: Huff, Philip, et al.
Veröffentlicht: (2026)
Phishing Detection System: An Ensemble Approach Using Character-Level CNN and Feature Engineering
von: Dubey, Rudra, et al.
Veröffentlicht: (2025)
von: Dubey, Rudra, et al.
Veröffentlicht: (2025)
Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions
von: Ma, Jianan, et al.
Veröffentlicht: (2026)
von: Ma, Jianan, et al.
Veröffentlicht: (2026)
Closing the Distribution Gap in Adversarial Training for LLMs
von: Hu, Chengzhi, et al.
Veröffentlicht: (2026)
von: Hu, Chengzhi, et al.
Veröffentlicht: (2026)
Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
von: Sanna, Arun Chowdary
Veröffentlicht: (2025)
von: Sanna, Arun Chowdary
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Hiremath Early Detection (HED) Score: A Measure-Theoretic Evaluation Standard for Temporal Intelligence
von: Hiremath, Prakul Sunil
Veröffentlicht: (2026) -
A High-Recall Cost-Sensitive Machine Learning Framework for Real-Time Online Banking Transaction Fraud Detection
von: R., Karthikeyan V., et al.
Veröffentlicht: (2026) -
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
von: Chakraborty, Amit, et al.
Veröffentlicht: (2025) -
Multi-Agent Honeypot-Based Request-Response Context Dataset for Improved SQL Injection Detection Performance
von: Yu, Hao, et al.
Veröffentlicht: (2026) -
Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI
von: Hoang, Tien Dat
Veröffentlicht: (2025)