GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Rad, Melissa Kazemi, Purpura, Alberto, Kumar, Himanshu, Chen, Emily, Sorower, Mohammad Shahed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
Supporting Human Raters with the Detection of Harmful Content using Large Language Models
by: Thomas, Kurt, et al.
Published: (2024)
by: Thomas, Kurt, et al.
Published: (2024)
Contrastive-KAN: A Semi-Supervised Intrusion Detection Framework for Cybersecurity with scarce Labeled Data
by: Alikhani, Mohammad, et al.
Published: (2025)
by: Alikhani, Mohammad, et al.
Published: (2025)
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
by: Yang, Jirui, et al.
Published: (2025)
by: Yang, Jirui, et al.
Published: (2025)
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
by: Zhang, Zelin, et al.
Published: (2026)
by: Zhang, Zelin, et al.
Published: (2026)
Secure Cross-Silo Synthetic Genomic Data Generation
by: Filienko, Daniil, et al.
Published: (2026)
by: Filienko, Daniil, et al.
Published: (2026)
Advancing CAN Network Security through RBM-Based Synthetic Attack Data Generation for Intrusion Detection Systems
by: Li, Huacheng, et al.
Published: (2025)
by: Li, Huacheng, et al.
Published: (2025)
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
by: Liu, Yuexiao, et al.
Published: (2025)
by: Liu, Yuexiao, et al.
Published: (2025)
Improving Google A2A Protocol: Protecting Sensitive Data and Mitigating Unintended Harms in Multi-Agent Systems
by: Louck, Yedidel, et al.
Published: (2025)
by: Louck, Yedidel, et al.
Published: (2025)
Advancing Vulnerability Classification with BERT: A Multi-Objective Learning Model
by: Tiwari, Himanshu
Published: (2025)
by: Tiwari, Himanshu
Published: (2025)
Securing Generative AI Agentic Workflows: Risks, Mitigation, and a Proposed Firewall Architecture
by: Bahadur, Sunil Kumar Jang, et al.
Published: (2025)
by: Bahadur, Sunil Kumar Jang, et al.
Published: (2025)
FHAIM: Fully Homomorphic AIM For Private Synthetic Data Generation
by: Kumar, Mayank, et al.
Published: (2026)
by: Kumar, Mayank, et al.
Published: (2026)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models
by: Purpura, Alberto, et al.
Published: (2025)
by: Purpura, Alberto, et al.
Published: (2025)
HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
by: Narula, Sidhant, et al.
Published: (2025)
by: Narula, Sidhant, et al.
Published: (2025)
CaPS: Collaborative and Private Synthetic Data Generation from Distributed Sources
by: Pentyala, Sikha, et al.
Published: (2024)
by: Pentyala, Sikha, et al.
Published: (2024)
Generative Models for Synthetic Urban Mobility Data: A Systematic Literature Review
by: Kapp, Alexandra, et al.
Published: (2024)
by: Kapp, Alexandra, et al.
Published: (2024)
A Comparison of SynDiffix Multi-table versus Single-table Synthetic Data
by: Francis, Paul
Published: (2024)
by: Francis, Paul
Published: (2024)
Red-MIRROR: Agentic LLM-based Autonomous Penetration Testing with Reflective Verification and Knowledge-augmented Interaction
by: Khang, Tran Vy, et al.
Published: (2026)
by: Khang, Tran Vy, et al.
Published: (2026)
Unbundle-Rewrite-Rebundle: Runtime Detection and Rewriting of Privacy-Harming Code in JavaScript Bundles
by: Ali, Mir Masood, et al.
Published: (2024)
by: Ali, Mir Masood, et al.
Published: (2024)
Scaling While Privacy Preserving: A Comprehensive Synthetic Tabular Data Generation and Evaluation in Learning Analytics
by: Liu, Qinyi, et al.
Published: (2024)
by: Liu, Qinyi, et al.
Published: (2024)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic Data
by: Zeng, Shenglai, et al.
Published: (2024)
by: Zeng, Shenglai, et al.
Published: (2024)
PCEvolve: Private Contrastive Evolution for Synthetic Dataset Generation via Few-Shot Private Data and Generative APIs
by: Zhang, Jianqing, et al.
Published: (2025)
by: Zhang, Jianqing, et al.
Published: (2025)
TableMark: A Multi-bit Watermark for Synthetic Tabular Data
by: Xia, Yuyang, et al.
Published: (2026)
by: Xia, Yuyang, et al.
Published: (2026)
Generative AI for Critical Infrastructure in Smart Grids: A Unified Framework for Synthetic Data Generation and Anomaly Detection
by: Zaboli, Aydin, et al.
Published: (2025)
by: Zaboli, Aydin, et al.
Published: (2025)
Enhancing Leakage Attacks on Searchable Symmetric Encryption Using LLM-Based Synthetic Data Generation
by: Chiu, Joshua, et al.
Published: (2025)
by: Chiu, Joshua, et al.
Published: (2025)
SynthGuard: Redefining Synthetic Data Generation with a Scalable and Privacy-Preserving Workflow Framework
by: Brito, Eduardo, et al.
Published: (2025)
by: Brito, Eduardo, et al.
Published: (2025)
Persistent Human Feedback, LLMs, and Static Analyzers for Secure Code Generation and Vulnerability Detection
by: Firouzi, Ehsan, et al.
Published: (2026)
by: Firouzi, Ehsan, et al.
Published: (2026)
End to End Collaborative Synthetic Data Generation
by: Pentyala, Sikha, et al.
Published: (2024)
by: Pentyala, Sikha, et al.
Published: (2024)
An Agentic Workflow for Detecting Personally Identifiable Information in Crash Narratives
by: Ma, Junyi, et al.
Published: (2026)
by: Ma, Junyi, et al.
Published: (2026)
Byzantine Failures Harm the Generalization of Robust Distributed Learning Algorithms More Than Data Poisoning
by: Boudou, Thomas, et al.
Published: (2025)
by: Boudou, Thomas, et al.
Published: (2025)
A Review of Privacy Metrics for Privacy-Preserving Synthetic Data Generation
by: Trudslev, Frederik Marinus, et al.
Published: (2025)
by: Trudslev, Frederik Marinus, et al.
Published: (2025)
KiNETGAN: Enabling Distributed Network Intrusion Detection through Knowledge-Infused Synthetic Data Generation
by: Kotal, Anantaa, et al.
Published: (2024)
by: Kotal, Anantaa, et al.
Published: (2024)
Efficient Jailbreaking of Large Models by Freeze Training: Lower Layers Exhibit Greater Sensitivity to Harmful Content
by: Shen, Hongyuan, et al.
Published: (2025)
by: Shen, Hongyuan, et al.
Published: (2025)
Demystifying and Detecting Agentic Workflow Injection Vulnerabilities in GitHub Actions
by: Wang, Shenao, et al.
Published: (2026)
by: Wang, Shenao, et al.
Published: (2026)
CLASP: Cost-Optimized LLM-based Agentic System for Phishing Detection
by: Trad, Fouad, et al.
Published: (2025)
by: Trad, Fouad, et al.
Published: (2025)
A Synthetic Conversational Smishing Dataset for Social Engineering Detection
by: Lochstampfor, Carl, et al.
Published: (2026)
by: Lochstampfor, Carl, et al.
Published: (2026)
CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation
by: Manuel, Dylan, et al.
Published: (2025)
by: Manuel, Dylan, et al.
Published: (2025)
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
by: Lu, Guoxin, et al.
Published: (2026)
by: Lu, Guoxin, et al.
Published: (2026)
Similar Items
-
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025) -
Supporting Human Raters with the Detection of Harmful Content using Large Language Models
by: Thomas, Kurt, et al.
Published: (2024) -
Contrastive-KAN: A Semi-Supervised Intrusion Detection Framework for Cybersecurity with scarce Labeled Data
by: Alikhani, Mohammad, et al.
Published: (2025) -
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
by: Yang, Jirui, et al.
Published: (2025) -
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
by: Zhang, Zelin, et al.
Published: (2026)