The Structural Safety Generalization Problem
Fuente:
arXiv
Salvato in:
| Autori principali: | Broomfield, Julius, Gibbs, Tom, Kosak-Hine, Ethan, Ingebretsen, George, Nasir, Tia, Zhang, Jason, Iranmanesh, Reihaneh, Pieri, Sara, Rabbany, Reihaneh, Pelrine, Kellin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
di: Gibbs, Tom, et al.
Pubblicazione: (2024)
di: Gibbs, Tom, et al.
Pubblicazione: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
di: Murphy, Brendan, et al.
Pubblicazione: (2025)
di: Murphy, Brendan, et al.
Pubblicazione: (2025)
Online Influence Campaigns: Strategies and Vulnerabilities
di: Musulan, Andreea, et al.
Pubblicazione: (2024)
di: Musulan, Andreea, et al.
Pubblicazione: (2024)
Hybrid Encryption with Certified Deletion in Preprocessing Model
di: Dey, Kunal, et al.
Pubblicazione: (2026)
di: Dey, Kunal, et al.
Pubblicazione: (2026)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
di: Struppek, Lukas, et al.
Pubblicazione: (2026)
di: Struppek, Lukas, et al.
Pubblicazione: (2026)
Robust and Reusable Fuzzy Extractors for Low-entropy Rate Randomness Sources
di: Panja, Somnath, et al.
Pubblicazione: (2024)
di: Panja, Somnath, et al.
Pubblicazione: (2024)
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
di: Rivera, Mauricio, et al.
Pubblicazione: (2024)
di: Rivera, Mauricio, et al.
Pubblicazione: (2024)
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
di: Vergho, Tyler, et al.
Pubblicazione: (2024)
di: Vergho, Tyler, et al.
Pubblicazione: (2024)
Secure Composition of Quantum Key Distribution and Symmetric Key Encryption
di: Dey, Kunal, et al.
Pubblicazione: (2025)
di: Dey, Kunal, et al.
Pubblicazione: (2025)
CCA-Secure Hybrid Encryption in Correlated Randomness Model and KEM Combiners
di: Panja, Somnath, et al.
Pubblicazione: (2024)
di: Panja, Somnath, et al.
Pubblicazione: (2024)
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
di: Hossain, Saad, et al.
Pubblicazione: (2026)
di: Hossain, Saad, et al.
Pubblicazione: (2026)
Scaling Trends for Data Poisoning in LLMs
di: Bowen, Dillon, et al.
Pubblicazione: (2024)
di: Bowen, Dillon, et al.
Pubblicazione: (2024)
Not What You Asked For: Typographic Attacks in Household Robot Manipulation
di: Iranmanesh, Ali, et al.
Pubblicazione: (2026)
di: Iranmanesh, Ali, et al.
Pubblicazione: (2026)
Exploiting Novel GPT-4 APIs
di: Pelrine, Kellin, et al.
Pubblicazione: (2023)
di: Pelrine, Kellin, et al.
Pubblicazione: (2023)
MAGIQ: A Post-Quantum Multi-Agentic AI Governance System with Provable Security
di: Avizheh, Sepideh, et al.
Pubblicazione: (2026)
di: Avizheh, Sepideh, et al.
Pubblicazione: (2026)
Uncertainty Resolution in Misinformation Detection
di: Orlovskiy, Yury, et al.
Pubblicazione: (2024)
di: Orlovskiy, Yury, et al.
Pubblicazione: (2024)
Accountable authentication with privacy protection: The Larch system for universal login
di: Dauterman, Emma, et al.
Pubblicazione: (2023)
di: Dauterman, Emma, et al.
Pubblicazione: (2023)
Generalized Implicit Factorization Problem
di: Feng, Yansong, et al.
Pubblicazione: (2023)
di: Feng, Yansong, et al.
Pubblicazione: (2023)
Exploiting CPU Clock Modulation for Covert Communication Channel
di: Alam, Shariful, et al.
Pubblicazione: (2024)
di: Alam, Shariful, et al.
Pubblicazione: (2024)
PoseGuard: Pose-Guided Generation with Safety Guardrails
di: Wang, Kongxin, et al.
Pubblicazione: (2025)
di: Wang, Kongxin, et al.
Pubblicazione: (2025)
Matching Ranks Over Probability Yields Truly Deep Safety Alignment
di: Vega, Jason, et al.
Pubblicazione: (2025)
di: Vega, Jason, et al.
Pubblicazione: (2025)
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
Maybenot: A Framework for Traffic Analysis Defenses
di: Pulls, Tobias, et al.
Pubblicazione: (2023)
di: Pulls, Tobias, et al.
Pubblicazione: (2023)
Enabling High-Frequency Trading with Near-Instant, Trustless Cross-Chain Transactions via Pre-Signing Adaptor Signatures
di: Francolla, Ethan, et al.
Pubblicazione: (2025)
di: Francolla, Ethan, et al.
Pubblicazione: (2025)
AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents
di: Henke, Julius
Pubblicazione: (2025)
di: Henke, Julius
Pubblicazione: (2025)
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
FP-Agent: Fingerprinting AI Browsing Agents
di: Wang, Ethan, et al.
Pubblicazione: (2026)
di: Wang, Ethan, et al.
Pubblicazione: (2026)
Process-Mining of Hypertraces: Enabling Scalable Formal Security Verification of (Automotive) Network Architectures
di: Figge, Julius, et al.
Pubblicazione: (2026)
di: Figge, Julius, et al.
Pubblicazione: (2026)
zkStruDul: Programming zkSNARKs with Structural Duality
di: Krishnan, Rahul, et al.
Pubblicazione: (2025)
di: Krishnan, Rahul, et al.
Pubblicazione: (2025)
A Simulation System Towards Solving Societal-Scale Manipulation
di: Touzel, Maximilian Puelma, et al.
Pubblicazione: (2024)
di: Touzel, Maximilian Puelma, et al.
Pubblicazione: (2024)
Ethical Hacking and its role in Cybersecurity
di: Asif, Fatima, et al.
Pubblicazione: (2024)
di: Asif, Fatima, et al.
Pubblicazione: (2024)
Weak Supervision for Real World Graphs
di: Nair, Pratheeksha, et al.
Pubblicazione: (2025)
di: Nair, Pratheeksha, et al.
Pubblicazione: (2025)
Nonmalleable Progress Leakage
di: Cecchetti, Ethan
Pubblicazione: (2025)
di: Cecchetti, Ethan
Pubblicazione: (2025)
Behavioral Authentication for Security and Safety
di: Wang, Cheng, et al.
Pubblicazione: (2023)
di: Wang, Cheng, et al.
Pubblicazione: (2023)
Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis
di: Xu, Zhenhao, et al.
Pubblicazione: (2026)
di: Xu, Zhenhao, et al.
Pubblicazione: (2026)
Generating Text from Uniform Meaning Representation
di: Markle, Emma, et al.
Pubblicazione: (2025)
di: Markle, Emma, et al.
Pubblicazione: (2025)
The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents
di: Wu, Baoyuan, et al.
Pubblicazione: (2026)
di: Wu, Baoyuan, et al.
Pubblicazione: (2026)
The Kubernetes Security Landscape: AI-Driven Insights from Developer Discussions
di: Curtis, J. Alexander, et al.
Pubblicazione: (2024)
di: Curtis, J. Alexander, et al.
Pubblicazione: (2024)
Generate "Normal", Edit Poisoned: Branding Injection via Hint Embedding in Image Editing
di: Sun, Desen, et al.
Pubblicazione: (2026)
di: Sun, Desen, et al.
Pubblicazione: (2026)
Rethinking Fraud Safety Evaluation: Multi-Round Attacks Reveal Safety-Utility Tradeoffs in Graph-Context LLM Defenders
di: Jiang, Laura, et al.
Pubblicazione: (2026)
di: Jiang, Laura, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
di: Gibbs, Tom, et al.
Pubblicazione: (2024) -
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
di: Murphy, Brendan, et al.
Pubblicazione: (2025) -
Online Influence Campaigns: Strategies and Vulnerabilities
di: Musulan, Andreea, et al.
Pubblicazione: (2024) -
Hybrid Encryption with Certified Deletion in Preprocessing Model
di: Dey, Kunal, et al.
Pubblicazione: (2026) -
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
di: Struppek, Lukas, et al.
Pubblicazione: (2026)