SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Hanxiu, Zheng, Yue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
por: Chen, Sixu, et al.
Publicado: (2026)
por: Chen, Sixu, et al.
Publicado: (2026)
BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
por: Gill, Waris, et al.
Publicado: (2025)
por: Gill, Waris, et al.
Publicado: (2025)
Instructional Fingerprinting of Large Language Models
por: Xu, Jiashu, et al.
Publicado: (2024)
por: Xu, Jiashu, et al.
Publicado: (2024)
Are Robust LLM Fingerprints Adversarially Robust?
por: Nasery, Anshul, et al.
Publicado: (2025)
por: Nasery, Anshul, et al.
Publicado: (2025)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
por: Dobre, David, et al.
Publicado: (2025)
por: Dobre, David, et al.
Publicado: (2025)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
por: Cao, Bochuan, et al.
Publicado: (2023)
por: Cao, Bochuan, et al.
Publicado: (2023)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
por: Labiad, Ismail, et al.
Publicado: (2025)
por: Labiad, Ismail, et al.
Publicado: (2025)
Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses
por: Ahmed, Mohamed, et al.
Publicado: (2025)
por: Ahmed, Mohamed, et al.
Publicado: (2025)
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
por: Alhazbi, Saeif, et al.
Publicado: (2025)
por: Alhazbi, Saeif, et al.
Publicado: (2025)
Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates
por: Zheng, Xiaosen, et al.
Publicado: (2024)
por: Zheng, Xiaosen, et al.
Publicado: (2024)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
por: Wang, Kun, et al.
Publicado: (2025)
por: Wang, Kun, et al.
Publicado: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
por: Wang, Yifei, et al.
Publicado: (2024)
por: Wang, Yifei, et al.
Publicado: (2024)
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
por: Zhang, Haoyu, et al.
Publicado: (2026)
por: Zhang, Haoyu, et al.
Publicado: (2026)
SVIP: Towards Verifiable Inference of Open-source Large Language Models
por: Sun, Yifan, et al.
Publicado: (2024)
por: Sun, Yifan, et al.
Publicado: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
por: Wu, Tong, et al.
Publicado: (2024)
por: Wu, Tong, et al.
Publicado: (2024)
PostMark: A Robust Blackbox Watermark for Large Language Models
por: Chang, Yapei, et al.
Publicado: (2024)
por: Chang, Yapei, et al.
Publicado: (2024)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
por: Yuan, Hongbang, et al.
Publicado: (2024)
por: Yuan, Hongbang, et al.
Publicado: (2024)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
por: Li, Yuexin, et al.
Publicado: (2026)
por: Li, Yuexin, et al.
Publicado: (2026)
From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
por: Xu, Xiaoyu, et al.
Publicado: (2026)
por: Xu, Xiaoyu, et al.
Publicado: (2026)
Linking Cryptoasset Attribution Tags to Knowledge Graph Entities: An LLM-based Approach
por: Avice, Régnier, et al.
Publicado: (2025)
por: Avice, Régnier, et al.
Publicado: (2025)
Towards Understanding the Robustness of Sparse Autoencoders
por: Saiyed, Ahson, et al.
Publicado: (2026)
por: Saiyed, Ahson, et al.
Publicado: (2026)
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
por: Zhang, Mohan, et al.
Publicado: (2026)
por: Zhang, Mohan, et al.
Publicado: (2026)
Federated In-Context LLM Agent Learning
por: Wu, Panlong, et al.
Publicado: (2024)
por: Wu, Panlong, et al.
Publicado: (2024)
STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents
por: Li, Jing-Jing, et al.
Publicado: (2025)
por: Li, Jing-Jing, et al.
Publicado: (2025)
Towards Building a Robust Toxicity Predictor
por: Bespalov, Dmitriy, et al.
Publicado: (2024)
por: Bespalov, Dmitriy, et al.
Publicado: (2024)
On Adversarial Robustness of Language Models in Transfer Learning
por: Turbal, Bohdan, et al.
Publicado: (2024)
por: Turbal, Bohdan, et al.
Publicado: (2024)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
por: Zhu, Sicheng, et al.
Publicado: (2024)
por: Zhu, Sicheng, et al.
Publicado: (2024)
Policy-Invisible Violations in LLM-Based Agents
por: Wu, Jie, et al.
Publicado: (2026)
por: Wu, Jie, et al.
Publicado: (2026)
Certifying LLM Safety against Adversarial Prompting
por: Kumar, Aounon, et al.
Publicado: (2023)
por: Kumar, Aounon, et al.
Publicado: (2023)
Directional Embedding Smoothing for Robust Vision Language Models
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
por: Xu, Xiaoyu, et al.
Publicado: (2025)
por: Xu, Xiaoyu, et al.
Publicado: (2025)
GaussMark: A Practical Approach for Structural Watermarking of Language Models
por: Block, Adam, et al.
Publicado: (2025)
por: Block, Adam, et al.
Publicado: (2025)
Adversarial Text Purification: A Large Language Model Approach for Defense
por: Moraffah, Raha, et al.
Publicado: (2024)
por: Moraffah, Raha, et al.
Publicado: (2024)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
por: Chrabąszcz, Maciej, et al.
Publicado: (2025)
por: Chrabąszcz, Maciej, et al.
Publicado: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
por: Halawi, Danny, et al.
Publicado: (2024)
por: Halawi, Danny, et al.
Publicado: (2024)
Adaptive Instruction Composition for Automated LLM Red-Teaming
por: Zymet, Jesse, et al.
Publicado: (2026)
por: Zymet, Jesse, et al.
Publicado: (2026)
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
por: Xu, Xiaoyu, et al.
Publicado: (2025)
por: Xu, Xiaoyu, et al.
Publicado: (2025)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
por: Gu, Tianle, et al.
Publicado: (2025)
por: Gu, Tianle, et al.
Publicado: (2025)
FIT to Forget: Robust Continual Unlearning for Large Language Models
por: Xu, Xiaoyu, et al.
Publicado: (2026)
por: Xu, Xiaoyu, et al.
Publicado: (2026)
Attacks and Defenses Against LLM Fingerprinting
por: Kurian, Kevin, et al.
Publicado: (2025)
por: Kurian, Kevin, et al.
Publicado: (2025)
Ejemplares similares
-
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
por: Chen, Sixu, et al.
Publicado: (2026) -
BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
por: Gill, Waris, et al.
Publicado: (2025) -
Instructional Fingerprinting of Large Language Models
por: Xu, Jiashu, et al.
Publicado: (2024) -
Are Robust LLM Fingerprints Adversarially Robust?
por: Nasery, Anshul, et al.
Publicado: (2025) -
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
por: Dobre, David, et al.
Publicado: (2025)