Training AI to be Loyal
Fuente:
arXiv
Guardado en:
| Autores principales: | Oh, Sewoong, Tyagi, Himanshu, Viswanath, Pramod |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Independence Criterion in Machine Unlearning of Features and Labels
por: Han, Ling, et al.
Publicado: (2024)
por: Han, Ling, et al.
Publicado: (2024)
A Novel Self-Attention-Enabled Weighted Ensemble-Based Convolutional Neural Network Framework for Distributed Denial of Service Attack Classification
por: S, Kanthimathi, et al.
Publicado: (2024)
por: S, Kanthimathi, et al.
Publicado: (2024)
Closing the Distribution Gap in Adversarial Training for LLMs
por: Hu, Chengzhi, et al.
Publicado: (2026)
por: Hu, Chengzhi, et al.
Publicado: (2026)
Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
por: Sanna, Arun Chowdary
Publicado: (2025)
por: Sanna, Arun Chowdary
Publicado: (2025)
Unlearning at Scale: Implementing the Right to be Forgotten in Large Language Models
por: X, Abdullah
Publicado: (2025)
por: X, Abdullah
Publicado: (2025)
MathLedger: A Verifiable Learning Substrate with Ledger-Attested Feedback
por: Abdullah, Ismail Ahmad
Publicado: (2025)
por: Abdullah, Ismail Ahmad
Publicado: (2025)
SAND: A Self-supervised and Adaptive NAS-Driven Framework for Hardware Trojan Detection
por: Pan, Zhixin, et al.
Publicado: (2025)
por: Pan, Zhixin, et al.
Publicado: (2025)
Evaluating Query Efficiency and Accuracy of Transfer Learning-based Model Extraction Attack in Federated Learning
por: Ahamed, Sayyed Farid, et al.
Publicado: (2025)
por: Ahamed, Sayyed Farid, et al.
Publicado: (2025)
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures
por: Kumar, Sachin
Publicado: (2026)
por: Kumar, Sachin
Publicado: (2026)
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity
por: Rashidi, Mohammadreza
Publicado: (2026)
por: Rashidi, Mohammadreza
Publicado: (2026)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
por: Gowda, Ishrith
Publicado: (2026)
por: Gowda, Ishrith
Publicado: (2026)
SCAFDS: Edge-Feature Graph Attention for Interbank Fraud Detection with Attribution-Grounded SAR Generation
por: Uddin, Mohammad Nasir
Publicado: (2026)
por: Uddin, Mohammad Nasir
Publicado: (2026)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
por: Lelle, Travis
Publicado: (2026)
por: Lelle, Travis
Publicado: (2026)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
por: Seddik, Issam, et al.
Publicado: (2025)
por: Seddik, Issam, et al.
Publicado: (2025)
XFED: Non-Collusive Model Poisoning Attack Against Byzantine-Robust Federated Classifiers
por: Mouri, Israt Jahan, et al.
Publicado: (2026)
por: Mouri, Israt Jahan, et al.
Publicado: (2026)
Attacking interpretable NLP systems
por: Abdukhamidov, Eldor, et al.
Publicado: (2025)
por: Abdukhamidov, Eldor, et al.
Publicado: (2025)
Inverting Cryptographic Hash Functions via Cube-and-Conquer
por: Zaikin, Oleg
Publicado: (2022)
por: Zaikin, Oleg
Publicado: (2022)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
por: DeLeeuw, Caleb
Publicado: (2026)
por: DeLeeuw, Caleb
Publicado: (2026)
David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning
por: Nellessen, Samuel, et al.
Publicado: (2026)
por: Nellessen, Samuel, et al.
Publicado: (2026)
Signal-Based Malware Classification Using 1D CNNs
por: Wilkie, Jack, et al.
Publicado: (2025)
por: Wilkie, Jack, et al.
Publicado: (2025)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
por: Nathanson, Samuel, et al.
Publicado: (2025)
por: Nathanson, Samuel, et al.
Publicado: (2025)
Are Robust LLM Fingerprints Adversarially Robust?
por: Nasery, Anshul, et al.
Publicado: (2025)
por: Nasery, Anshul, et al.
Publicado: (2025)
Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI
por: Rath, Plawan Kumar, et al.
Publicado: (2026)
por: Rath, Plawan Kumar, et al.
Publicado: (2026)
Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents
por: Ravindran, Santhosh Kumar
Publicado: (2026)
por: Ravindran, Santhosh Kumar
Publicado: (2026)
FedAttr: Towards Privacy-preserving Client-Level Attribution in Federated LLM Fine-tuning
por: Zhang, Su, et al.
Publicado: (2026)
por: Zhang, Su, et al.
Publicado: (2026)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
por: Blanco-Justicia, Alberto, et al.
Publicado: (2024)
por: Blanco-Justicia, Alberto, et al.
Publicado: (2024)
Towards Low-Latency and Adaptive Ransomware Detection Using Contrastive Learning
por: Pan, Zhixin, et al.
Publicado: (2025)
por: Pan, Zhixin, et al.
Publicado: (2025)
Monotonicity as an Architectural Bias for Robust Language Models
por: Cooper, Patrick, et al.
Publicado: (2026)
por: Cooper, Patrick, et al.
Publicado: (2026)
Et Tu Certifications: Robustness Certificates Yield Better Adversarial Examples
por: Cullen, Andrew C., et al.
Publicado: (2023)
por: Cullen, Andrew C., et al.
Publicado: (2023)
Contrastive Self-Supervised Network Intrusion Detection using Augmented Negative Pairs
por: Wilkie, Jack, et al.
Publicado: (2025)
por: Wilkie, Jack, et al.
Publicado: (2025)
Few-Shot Network Intrusion Detection Using Online Triplet Mining
por: Wilkie, Jack, et al.
Publicado: (2026)
por: Wilkie, Jack, et al.
Publicado: (2026)
A Novel Contrastive Loss for Zero-Day Network Intrusion Detection
por: Wilkie, Jack, et al.
Publicado: (2026)
por: Wilkie, Jack, et al.
Publicado: (2026)
Improving the Convergence Rate of Ray Search Optimization for Query-Efficient Hard-Label Attacks
por: Xu, Xinjie, et al.
Publicado: (2025)
por: Xu, Xinjie, et al.
Publicado: (2025)
One-vs.-One Mitigation of Intersectional Bias: A General Method to Extend Fairness-Aware Binary Classification
por: Kobayashi, Kenji, et al.
Publicado: (2020)
por: Kobayashi, Kenji, et al.
Publicado: (2020)
Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning
por: Subedar, Noah, et al.
Publicado: (2025)
por: Subedar, Noah, et al.
Publicado: (2025)
Density-aware Sample-specific Attack
por: Wang, Qiyuan, et al.
Publicado: (2026)
por: Wang, Qiyuan, et al.
Publicado: (2026)
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction
por: Garcia, Gabriel
Publicado: (2026)
por: Garcia, Gabriel
Publicado: (2026)
Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
por: Cotti, Luca, et al.
Publicado: (2025)
por: Cotti, Luca, et al.
Publicado: (2025)
One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIP
por: Xu, Binyan, et al.
Publicado: (2025)
por: Xu, Binyan, et al.
Publicado: (2025)
Public-Decay Homomorphic State Space Models for Private Sequence Inference
por: Brito, Luis
Publicado: (2026)
por: Brito, Luis
Publicado: (2026)
Ejemplares similares
-
Towards Independence Criterion in Machine Unlearning of Features and Labels
por: Han, Ling, et al.
Publicado: (2024) -
A Novel Self-Attention-Enabled Weighted Ensemble-Based Convolutional Neural Network Framework for Distributed Denial of Service Attack Classification
por: S, Kanthimathi, et al.
Publicado: (2024) -
Closing the Distribution Gap in Adversarial Training for LLMs
por: Hu, Chengzhi, et al.
Publicado: (2026) -
Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
por: Sanna, Arun Chowdary
Publicado: (2025) -
Unlearning at Scale: Implementing the Right to be Forgotten in Large Language Models
por: X, Abdullah
Publicado: (2025)