Quantifying CBRN Risk in Frontier Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Divyanshu, Birur, Nitin Aravind, Baswa, Tanay, Agarwal, Sahil, Harshangi, Prashanth |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
No Free Lunch with Guardrails
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
by: Kumar, Anurakt, et al.
Published: (2024)
by: Kumar, Anurakt, et al.
Published: (2024)
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
by: Kumar, Divyanshu, et al.
Published: (2024)
by: Kumar, Divyanshu, et al.
Published: (2024)
VERA: Validation and Enhancement for Retrieval Augmented systems
by: Birur, Nitin Aravind, et al.
Published: (2024)
by: Birur, Nitin Aravind, et al.
Published: (2024)
SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models
by: Kumar, Divyanshu, et al.
Published: (2026)
by: Kumar, Divyanshu, et al.
Published: (2026)
Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments
by: Kumar, Divyanshu, et al.
Published: (2026)
by: Kumar, Divyanshu, et al.
Published: (2026)
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
by: Kumar, Divyanshu, et al.
Published: (2024)
by: Kumar, Divyanshu, et al.
Published: (2024)
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
by: Marchand, Rahul, et al.
Published: (2026)
by: Marchand, Rahul, et al.
Published: (2026)
AgentWall: A Runtime Safety Layer for Local AI Agents
by: Aravind, Ashwin
Published: (2026)
by: Aravind, Ashwin
Published: (2026)
Benchmarking Large Language Models for Zero-shot and Few-shot Phishing URL Detection
by: Hasan, Najmul, et al.
Published: (2026)
by: Hasan, Najmul, et al.
Published: (2026)
Automated Classification of Cybercrime Complaints using Transformer-based Language Models for Hinglish Texts
by: Rani, Nanda, et al.
Published: (2024)
by: Rani, Nanda, et al.
Published: (2024)
Precision Guided Approach to Mitigate Data Poisoning Attacks in Federated Learning
by: Kumar, K Naveen, et al.
Published: (2024)
by: Kumar, K Naveen, et al.
Published: (2024)
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
by: Li, Changyi, et al.
Published: (2026)
by: Li, Changyi, et al.
Published: (2026)
Quantum Gatekeeper: Multi-Factor Context-Bound Image Steganography with VQC Based Key Derivation on Quantum Hardware
by: Tomar, Sahil, et al.
Published: (2026)
by: Tomar, Sahil, et al.
Published: (2026)
From Incomplete Architecture to Quantified Risk: Multimodal LLM-Driven Security Assessment for Cyber-Physical Systems
by: Huang, Shaofei, et al.
Published: (2026)
by: Huang, Shaofei, et al.
Published: (2026)
A Survey on Offensive AI Within Cybersecurity
by: Girhepuje, Sahil, et al.
Published: (2024)
by: Girhepuje, Sahil, et al.
Published: (2024)
Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks
by: Dahiya, Vivek, et al.
Published: (2026)
by: Dahiya, Vivek, et al.
Published: (2026)
Mythos and the Unverified Cage: Z3-Based Pre-Deployment Verification for Frontier-Model Sandbox Infrastructure
by: Blain, Dominik
Published: (2026)
by: Blain, Dominik
Published: (2026)
SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence
by: Aghaei, Ehsan, et al.
Published: (2025)
by: Aghaei, Ehsan, et al.
Published: (2025)
Jailbroken Frontier Models Retain Their Capabilities
by: Zhu, Daniel, et al.
Published: (2026)
by: Zhu, Daniel, et al.
Published: (2026)
LLM Embedding-based Attribution (LEA): Quantifying Source Contributions to Generative Model's Response for Vulnerability Analysis
by: Fayyazi, Reza, et al.
Published: (2025)
by: Fayyazi, Reza, et al.
Published: (2025)
SecureRAG-RTL: A Retrieval-Augmented, Multi-Agent, Zero-Shot LLM-Driven Framework for Hardware Vulnerability Detection
by: Hasan, Touseef, et al.
Published: (2026)
by: Hasan, Touseef, et al.
Published: (2026)
Quantifying Loss Aversion in Cyber Adversaries via LLM Analysis
by: Hans, Soham, et al.
Published: (2025)
by: Hans, Soham, et al.
Published: (2025)
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
by: Tong, Haibo, et al.
Published: (2026)
by: Tong, Haibo, et al.
Published: (2026)
Quantifying and Defending against Privacy Threats on Federated Knowledge Graph Embedding
by: Hu, Yuke, et al.
Published: (2023)
by: Hu, Yuke, et al.
Published: (2023)
Jailbreaking Frontier Foundation Models Through Intention Deception
by: Wang, Xinhe, et al.
Published: (2026)
by: Wang, Xinhe, et al.
Published: (2026)
Internal Safety Collapse in Frontier Large Language Models
by: Wu, Yutao, et al.
Published: (2026)
by: Wu, Yutao, et al.
Published: (2026)
Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis
by: Topol, Zvi
Published: (2026)
by: Topol, Zvi
Published: (2026)
Quantifying AI Vulnerabilities: A Synthesis of Complexity, Dynamical Systems, and Game Theory
by: Kereopa-Yorke, B
Published: (2024)
by: Kereopa-Yorke, B
Published: (2024)
MCP-in-SoS: Risk assessment framework for open-source MCP servers
by: Kumar, Pratyay, et al.
Published: (2026)
by: Kumar, Pratyay, et al.
Published: (2026)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
by: Gibbs, Tom, et al.
Published: (2024)
by: Gibbs, Tom, et al.
Published: (2024)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
New Wide-Net-Casting Jailbreak Attacks Risk Large Models
by: Xiang, Qiuchi, et al.
Published: (2026)
by: Xiang, Qiuchi, et al.
Published: (2026)
Quantifying Security Vulnerabilities: A Metric-Driven Security Analysis of Gaps in Current AI Standards
by: Madhavan, Keerthana, et al.
Published: (2025)
by: Madhavan, Keerthana, et al.
Published: (2025)
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
by: Udeshi, Meet, et al.
Published: (2025)
by: Udeshi, Meet, et al.
Published: (2025)
RiskSEA : A Scalable Graph Embedding for Detecting On-chain Fraudulent Activities on the Ethereum Blockchain
by: Agarwal, Ayush, et al.
Published: (2024)
by: Agarwal, Ayush, et al.
Published: (2024)
Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
by: Li, Yakai, et al.
Published: (2025)
by: Li, Yakai, et al.
Published: (2025)
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
by: Teng, Ma, et al.
Published: (2024)
by: Teng, Ma, et al.
Published: (2024)
Similar Items
-
No Free Lunch with Guardrails
by: Kumar, Divyanshu, et al.
Published: (2025) -
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
by: Kumar, Anurakt, et al.
Published: (2024) -
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
by: Kumar, Divyanshu, et al.
Published: (2025) -
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
by: Kumar, Divyanshu, et al.
Published: (2024) -
VERA: Validation and Enhancement for Retrieval Augmented systems
by: Birur, Nitin Aravind, et al.
Published: (2024)