Instructional Fingerprinting of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Jiashu, Wang, Fei, Ma, Mingyu Derek, Koh, Pang Wei, Xiao, Chaowei, Chen, Muhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
von: Xu, Jiashu, et al.
Veröffentlicht: (2023)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
von: Alhazbi, Saeif, et al.
Veröffentlicht: (2025)
von: Alhazbi, Saeif, et al.
Veröffentlicht: (2025)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
On the Role of Attention Heads in Large Language Model Safety
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
ThinkGuard: Deliberative Slow Thinking Leads to Cautious Guardrails
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wen, Xiaofei, et al.
Veröffentlicht: (2025)
FIT to Forget: Robust Continual Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2026)
Duwak: Dual Watermarks in Large Language Models
von: Zhu, Chaoyi, et al.
Veröffentlicht: (2024)
von: Zhu, Chaoyi, et al.
Veröffentlicht: (2024)
A Study of Backdoors in Instruction Fine-tuned Language Models
von: Raghuram, Jayaram, et al.
Veröffentlicht: (2024)
von: Raghuram, Jayaram, et al.
Veröffentlicht: (2024)
EnJa: Ensemble Jailbreak on Large Language Models
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyu, et al.
Veröffentlicht: (2025)
Detecting Training Data of Large Language Models via Expectation Maximization
von: Kim, Gyuwan, et al.
Veröffentlicht: (2024)
von: Kim, Gyuwan, et al.
Veröffentlicht: (2024)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
von: Li, Xiao, et al.
Veröffentlicht: (2024)
von: Li, Xiao, et al.
Veröffentlicht: (2024)
Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
von: Wang, Zhepeng, et al.
Veröffentlicht: (2024)
von: Wang, Zhepeng, et al.
Veröffentlicht: (2024)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
von: Chen, Xiaoyi, et al.
Veröffentlicht: (2023)
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
von: Luo, Weidi, et al.
Veröffentlicht: (2024)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
Lifelong Safety Alignment for Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
von: Bai, Yang, et al.
Veröffentlicht: (2024)
von: Bai, Yang, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
Large Language Models in Cybersecurity: State-of-the-Art
von: Motlagh, Farzad Nourmohammadzadeh, et al.
Veröffentlicht: (2024)
von: Motlagh, Farzad Nourmohammadzadeh, et al.
Veröffentlicht: (2024)
Jailbreaking Large Language Models with Symbolic Mathematics
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
von: Bethany, Emet, et al.
Veröffentlicht: (2024)
Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting
von: Zhang, Hanxiu, et al.
Veröffentlicht: (2025)
von: Zhang, Hanxiu, et al.
Veröffentlicht: (2025)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
REEF: Representation Encoding Fingerprints for Large Language Models
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
von: Zhang, Jie, et al.
Veröffentlicht: (2024)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
Machine Unlearning of Pre-trained Large Language Models
von: Yao, Jin, et al.
Veröffentlicht: (2024)
von: Yao, Jin, et al.
Veröffentlicht: (2024)
Interpreting the Repeated Token Phenomenon in Large Language Models
von: Yona, Itay, et al.
Veröffentlicht: (2025)
von: Yona, Itay, et al.
Veröffentlicht: (2025)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
von: Zhao, Andrew, et al.
Veröffentlicht: (2024)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
von: Boreiko, Valentyn, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
von: Xu, Jiashu, et al.
Veröffentlicht: (2023) -
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
von: Liu, Qin, et al.
Veröffentlicht: (2024) -
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024) -
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023) -
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2023)