Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
Fuente:
arXiv
Salvato in:
| Autori principali: | Rosenblatt, Lucas, Liu, Peihan, McKenna, Ryan, Ponomareva, Natalia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
di: Leonesi, Matteo, et al.
Pubblicazione: (2026)
di: Leonesi, Matteo, et al.
Pubblicazione: (2026)
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
di: Young, Richard J., et al.
Pubblicazione: (2026)
di: Young, Richard J., et al.
Pubblicazione: (2026)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
di: Lucas, Tom, et al.
Pubblicazione: (2026)
di: Lucas, Tom, et al.
Pubblicazione: (2026)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
di: DeLeeuw, Caleb
Pubblicazione: (2026)
di: DeLeeuw, Caleb
Pubblicazione: (2026)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
di: Othman, Refat
Pubblicazione: (2026)
di: Othman, Refat
Pubblicazione: (2026)
AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
di: Nathanson, Samuel, et al.
Pubblicazione: (2025)
di: Nathanson, Samuel, et al.
Pubblicazione: (2025)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
di: Cohen, Liran, et al.
Pubblicazione: (2025)
di: Cohen, Liran, et al.
Pubblicazione: (2025)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
di: Bercovich, Ivan, et al.
Pubblicazione: (2026)
di: Bercovich, Ivan, et al.
Pubblicazione: (2026)
RADEP: A Resilient Adaptive Defense Framework Against Model Extraction Attacks
di: Chakraborty, Amit, et al.
Pubblicazione: (2025)
di: Chakraborty, Amit, et al.
Pubblicazione: (2025)
Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents
di: Sidik, Bronislav, et al.
Pubblicazione: (2026)
di: Sidik, Bronislav, et al.
Pubblicazione: (2026)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
di: Blanco-Justicia, Alberto, et al.
Pubblicazione: (2024)
di: Blanco-Justicia, Alberto, et al.
Pubblicazione: (2024)
Comparing Fairness of Generative Mobility Models
di: Wang, Daniel, et al.
Pubblicazione: (2024)
di: Wang, Daniel, et al.
Pubblicazione: (2024)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
di: Beltoft, Stine, et al.
Pubblicazione: (2025)
di: Beltoft, Stine, et al.
Pubblicazione: (2025)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
di: Yeste, Víctor, et al.
Pubblicazione: (2026)
di: Yeste, Víctor, et al.
Pubblicazione: (2026)
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
di: Dawson, Ads, et al.
Pubblicazione: (2025)
di: Dawson, Ads, et al.
Pubblicazione: (2025)
The Automation Advantage in AI Red Teaming
di: Mulla, Rob, et al.
Pubblicazione: (2025)
di: Mulla, Rob, et al.
Pubblicazione: (2025)
MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents
di: Gowda, Ishrith
Pubblicazione: (2026)
di: Gowda, Ishrith
Pubblicazione: (2026)
Scalable and Verifiable Federated Learning for Cross-Institution Financial Fraud Detection
di: Panth, Prajwal, et al.
Pubblicazione: (2026)
di: Panth, Prajwal, et al.
Pubblicazione: (2026)
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
di: Bommarito II, Michael J.
Pubblicazione: (2025)
di: Bommarito II, Michael J.
Pubblicazione: (2025)
Attacking interpretable NLP systems
di: Abdukhamidov, Eldor, et al.
Pubblicazione: (2025)
di: Abdukhamidov, Eldor, et al.
Pubblicazione: (2025)
Dr. Jekyll and Mr. Hyde: Two Faces of LLMs
di: Collu, Matteo Gioele, et al.
Pubblicazione: (2023)
di: Collu, Matteo Gioele, et al.
Pubblicazione: (2023)
AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification
di: Zhang, Tian, et al.
Pubblicazione: (2026)
di: Zhang, Tian, et al.
Pubblicazione: (2026)
Security Considerations for Multi-agent Systems
di: Nguyen, Tam, et al.
Pubblicazione: (2026)
di: Nguyen, Tam, et al.
Pubblicazione: (2026)
Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice
di: Ge, Yuxu
Pubblicazione: (2026)
di: Ge, Yuxu
Pubblicazione: (2026)
Sensitivity Uncertainty Alignment in Large Language Models
di: Hiremath, Prakul Sunil, et al.
Pubblicazione: (2026)
di: Hiremath, Prakul Sunil, et al.
Pubblicazione: (2026)
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
di: Li, Shenghao
Pubblicazione: (2025)
di: Li, Shenghao
Pubblicazione: (2025)
How Worrying Are Privacy Attacks Against Machine Learning?
di: Domingo-Ferrer, Josep
Pubblicazione: (2025)
di: Domingo-Ferrer, Josep
Pubblicazione: (2025)
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
di: Yeste, Víctor, et al.
Pubblicazione: (2026)
di: Yeste, Víctor, et al.
Pubblicazione: (2026)
Do Schwartz Higher-Order Values Help Sentence-Level Human Value Detection? A Study of Hierarchical Gating and Calibration
di: Yeste, Víctor, et al.
Pubblicazione: (2026)
di: Yeste, Víctor, et al.
Pubblicazione: (2026)
$\mathsf{OPA}$: One-shot Private Aggregation with Single Client Interaction and its Applications to Federated Learning
di: Karthikeyan, Harish, et al.
Pubblicazione: (2024)
di: Karthikeyan, Harish, et al.
Pubblicazione: (2024)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
di: Grofsky, Matthew
Pubblicazione: (2025)
di: Grofsky, Matthew
Pubblicazione: (2025)
Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study
di: Khatiwala, Jeel Piyushkumar, et al.
Pubblicazione: (2026)
di: Khatiwala, Jeel Piyushkumar, et al.
Pubblicazione: (2026)
Multilingual AI-Driven Password Strength Estimation with Similarity-Based Detection
di: Palaniappan, Nikitha M., et al.
Pubblicazione: (2026)
di: Palaniappan, Nikitha M., et al.
Pubblicazione: (2026)
Towards Agentic Investigation of Security Alerts
di: Eilertsen, Even, et al.
Pubblicazione: (2026)
di: Eilertsen, Even, et al.
Pubblicazione: (2026)
Measuring the Authority Stack of AI Systems: Empirical Analysis of 366,120 Forced-Choice Responses Across 8 AI Models
di: Lee, Seulki
Pubblicazione: (2026)
di: Lee, Seulki
Pubblicazione: (2026)
Who Governs the Machine? A Machine Identity Governance Taxonomy (MIGT) for AI Systems Operating Across Enterprise and Geopolitical Boundaries
di: Kurtz, Andrew, et al.
Pubblicazione: (2026)
di: Kurtz, Andrew, et al.
Pubblicazione: (2026)
Uncovering Bugs in Formal Explainers: A Case Study with PyXAI
di: Huang, Xuanxiang, et al.
Pubblicazione: (2025)
di: Huang, Xuanxiang, et al.
Pubblicazione: (2025)
Non-Adaptive Adversarial Face Generation
di: Kim, Sunpill, et al.
Pubblicazione: (2025)
di: Kim, Sunpill, et al.
Pubblicazione: (2025)
Toward Individual Fairness Without Centralized Data: Selective Counterfactual Consistency for Vertical Federated Learning
di: Wasif, Dawood, et al.
Pubblicazione: (2026)
di: Wasif, Dawood, et al.
Pubblicazione: (2026)
Can AI Keep a Secret? Contextual Integrity Verification: A Provable Security Architecture for LLMs
di: Gupta, Aayush
Pubblicazione: (2025)
di: Gupta, Aayush
Pubblicazione: (2025)
Documenti analoghi
-
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
di: Leonesi, Matteo, et al.
Pubblicazione: (2026) -
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025)
di: Young, Richard J., et al.
Pubblicazione: (2026) -
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
di: Lucas, Tom, et al.
Pubblicazione: (2026) -
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
di: DeLeeuw, Caleb
Pubblicazione: (2026) -
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
di: Othman, Refat
Pubblicazione: (2026)