Synthetic Data for Veterinary EHR De-identification: Benefits, Limits, and Safety Trade-offs Under Fixed Compute
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Brundage, David |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
par: Wang, Yanbo, et autres
Publié: (2026)
par: Wang, Yanbo, et autres
Publié: (2026)
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
par: Fonseca, Joao, et autres
Publié: (2025)
par: Fonseca, Joao, et autres
Publié: (2025)
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
par: Kumar, Anurakt, et autres
Publié: (2024)
par: Kumar, Anurakt, et autres
Publié: (2024)
Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation
par: Miranda, Michele, et autres
Publié: (2026)
par: Miranda, Michele, et autres
Publié: (2026)
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
par: Patel, Hitesh Laxmichand, et autres
Publié: (2024)
par: Patel, Hitesh Laxmichand, et autres
Publié: (2024)
Downstream Trade-offs of a Family of Text Watermarks
par: Ajith, Anirudh, et autres
Publié: (2023)
par: Ajith, Anirudh, et autres
Publié: (2023)
DataShield: Safety-degrading Data Filtering for LLM Benign Instruction Fine-Tuning
par: Zhang, Junbo, et autres
Publié: (2026)
par: Zhang, Junbo, et autres
Publié: (2026)
Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
par: Liang, Zi, et autres
Publié: (2025)
par: Liang, Zi, et autres
Publié: (2025)
What Matters For Safety Alignment?
par: Li, Xing, et autres
Publié: (2026)
par: Li, Xing, et autres
Publié: (2026)
An Independent Safety Evaluation of Kimi K2.5
par: Yong, Zheng-Xin, et autres
Publié: (2026)
par: Yong, Zheng-Xin, et autres
Publié: (2026)
Internal Safety Collapse in Frontier Large Language Models
par: Wu, Yutao, et autres
Publié: (2026)
par: Wu, Yutao, et autres
Publié: (2026)
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
par: Xu, Chejian, et autres
Publié: (2025)
par: Xu, Chejian, et autres
Publié: (2025)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
par: Liu, Songyang, et autres
Publié: (2025)
par: Liu, Songyang, et autres
Publié: (2025)
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
par: Zhou, Kaiwen, et autres
Publié: (2025)
par: Zhou, Kaiwen, et autres
Publié: (2025)
SGuard-v1: Safety Guardrail for Large Language Models
par: Lee, JoonHo, et autres
Publié: (2025)
par: Lee, JoonHo, et autres
Publié: (2025)
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
par: Lu, Guoxin, et autres
Publié: (2026)
par: Lu, Guoxin, et autres
Publié: (2026)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
par: Ying, Zonghao, et autres
Publié: (2025)
par: Ying, Zonghao, et autres
Publié: (2025)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
par: Wang, Zijun, et autres
Publié: (2026)
par: Wang, Zijun, et autres
Publié: (2026)
ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack
par: Li, Hao, et autres
Publié: (2026)
par: Li, Hao, et autres
Publié: (2026)
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
par: Schnabl, Christoph, et autres
Publié: (2025)
par: Schnabl, Christoph, et autres
Publié: (2025)
From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
par: Chae, Kyubyung, et autres
Publié: (2025)
par: Chae, Kyubyung, et autres
Publié: (2025)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
par: Xu, Zhangchen, et autres
Publié: (2024)
par: Xu, Zhangchen, et autres
Publié: (2024)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
par: He, Luxi, et autres
Publié: (2024)
par: He, Luxi, et autres
Publié: (2024)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
par: Wang, Kun, et autres
Publié: (2026)
par: Wang, Kun, et autres
Publié: (2026)
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
par: Nawal, Aditya, et autres
Publié: (2026)
par: Nawal, Aditya, et autres
Publié: (2026)
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
par: Gu, Tianle, et autres
Publié: (2024)
par: Gu, Tianle, et autres
Publié: (2024)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
par: Huang, Chao, et autres
Publié: (2025)
par: Huang, Chao, et autres
Publié: (2025)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
par: Zhao, Shuai, et autres
Publié: (2024)
par: Zhao, Shuai, et autres
Publié: (2024)
AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models
par: Zhang, Jinchuan, et autres
Publié: (2025)
par: Zhang, Jinchuan, et autres
Publié: (2025)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
par: Banayeeanzade, Amin, et autres
Publié: (2026)
par: Banayeeanzade, Amin, et autres
Publié: (2026)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
par: Li, Tianhao, et autres
Publié: (2024)
par: Li, Tianhao, et autres
Publié: (2024)
Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
par: Wang, Haoran, et autres
Publié: (2023)
par: Wang, Haoran, et autres
Publié: (2023)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
par: Fu, Tingchen, et autres
Publié: (2024)
par: Fu, Tingchen, et autres
Publié: (2024)
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
par: Su, Yanghao, et autres
Publié: (2026)
par: Su, Yanghao, et autres
Publié: (2026)
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
par: Leong, Chak Tou, et autres
Publié: (2025)
par: Leong, Chak Tou, et autres
Publié: (2025)
Dynamic Fog Computing for Enhanced LLM Execution in Medical Applications
par: Zagar, Philipp, et autres
Publié: (2024)
par: Zagar, Philipp, et autres
Publié: (2024)
DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts
par: Sun, Xiongtao, et autres
Publié: (2024)
par: Sun, Xiongtao, et autres
Publié: (2024)
When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents
par: Jones, Jaylen, et autres
Publié: (2026)
par: Jones, Jaylen, et autres
Publié: (2026)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
par: Wang, Kun, et autres
Publié: (2025)
par: Wang, Kun, et autres
Publié: (2025)
DeSIA: Attribute Inference Attacks Against Limited Fixed Aggregate Statistics
par: Mao, Yifeng, et autres
Publié: (2025)
par: Mao, Yifeng, et autres
Publié: (2025)
Documents similaires
-
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
par: Wang, Yanbo, et autres
Publié: (2026) -
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
par: Fonseca, Joao, et autres
Publié: (2025) -
SAGE-RT: Synthetic Alignment data Generation for Safety Evaluation and Red Teaming
par: Kumar, Anurakt, et autres
Publié: (2024) -
Differentially Private De-identification of Dutch Clinical Notes: A Comparative Evaluation
par: Miranda, Michele, et autres
Publié: (2026) -
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
par: Patel, Hitesh Laxmichand, et autres
Publié: (2024)