Intrinsic Model Weaknesses: How Priming Attacks Unveil Vulnerabilities in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yuyi, Zhan, Runzhe, Wong, Derek F., Chao, Lidia S., Tao, Ailin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
by: Huang, Yuyi, et al.
Published: (2025)
by: Huang, Yuyi, et al.
Published: (2025)
Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model
by: Xu, Haoyun, et al.
Published: (2024)
by: Xu, Haoyun, et al.
Published: (2024)
Rethinking Prompt-based Debiasing in Large Language Models
by: Yang, Xinyi, et al.
Published: (2025)
by: Yang, Xinyi, et al.
Published: (2025)
Prefix Text as a Yarn: Eliciting Non-English Alignment in Foundation Language Model
by: Zhan, Runzhe, et al.
Published: (2024)
by: Zhan, Runzhe, et al.
Published: (2024)
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
by: Zhan, Runzhe, et al.
Published: (2025)
by: Zhan, Runzhe, et al.
Published: (2025)
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
by: Chen, Xin, et al.
Published: (2026)
by: Chen, Xin, et al.
Published: (2026)
A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
by: Wu, Junchao, et al.
Published: (2023)
by: Wu, Junchao, et al.
Published: (2023)
Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation
by: Sun, Yanming, et al.
Published: (2025)
by: Sun, Yanming, et al.
Published: (2025)
Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
by: Wu, Junchao, et al.
Published: (2024)
by: Wu, Junchao, et al.
Published: (2024)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
by: Wu, Junchao, et al.
Published: (2024)
by: Wu, Junchao, et al.
Published: (2024)
Unveiling LLMs' Metaphorical Understanding: Exploring Conceptual Irrelevance, Context Leveraging and Syntactic Influence
by: Ye, Fengying, et al.
Published: (2025)
by: Ye, Fengying, et al.
Published: (2025)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
by: Ma, Jingkun, et al.
Published: (2024)
by: Ma, Jingkun, et al.
Published: (2024)
FOCUS: Forging Originality through Contrastive Use in Self-Plagiarism for Language Models
by: Lan, Kaixin, et al.
Published: (2024)
by: Lan, Kaixin, et al.
Published: (2024)
Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
by: Lian, Jiawei, et al.
Published: (2025)
by: Lian, Jiawei, et al.
Published: (2025)
Unveiling Linguistic Regions in Large Language Models
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
by: Dai, Runpeng, et al.
Published: (2025)
by: Dai, Runpeng, et al.
Published: (2025)
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
A Two-Stage Prediction-Aware Contrastive Learning Framework for Multi-Intent NLU
by: Chen, Guanhua, et al.
Published: (2024)
by: Chen, Guanhua, et al.
Published: (2024)
Weak-to-Strong Jailbreaking on Large Language Models
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination
by: He, Xiaoqi, et al.
Published: (2026)
by: He, Xiaoqi, et al.
Published: (2026)
Can ChatGPT Really Understand Modern Chinese Poetry?
by: Wang, Shanshan, et al.
Published: (2026)
by: Wang, Shanshan, et al.
Published: (2026)
What is the Best Way for ChatGPT to Translate Poetry?
by: Wang, Shanshan, et al.
Published: (2024)
by: Wang, Shanshan, et al.
Published: (2024)
Denial-of-Service Poisoning Attacks against Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
by: Zhang, Jiayi, et al.
Published: (2025)
by: Zhang, Jiayi, et al.
Published: (2025)
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
by: Bianchi, Federico, et al.
Published: (2024)
by: Bianchi, Federico, et al.
Published: (2024)
Unveiling Super Experts in Mixture-of-Experts Large Language Models
by: Su, Zunhai, et al.
Published: (2025)
by: Su, Zunhai, et al.
Published: (2025)
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
by: Ren, Juan, et al.
Published: (2025)
by: Ren, Juan, et al.
Published: (2025)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
SGIC: A Self-Guided Iterative Calibration Framework for RAG
by: Chen, Guanhua, et al.
Published: (2025)
by: Chen, Guanhua, et al.
Published: (2025)
Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs
by: Ford, Casey, et al.
Published: (2026)
by: Ford, Casey, et al.
Published: (2026)
Robustness of Large Language Models Against Adversarial Attacks
by: Tao, Yiyi, et al.
Published: (2024)
by: Tao, Yiyi, et al.
Published: (2024)
AuditWen:An Open-Source Large Language Model for Audit
by: Huang, Jiajia, et al.
Published: (2024)
by: Huang, Jiajia, et al.
Published: (2024)
Large Language Model for Multi-Domain Translation: Benchmarking and Domain CoT Fine-tuning
by: Hu, Tianxiang, et al.
Published: (2024)
by: Hu, Tianxiang, et al.
Published: (2024)
DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
by: Li, Zherui, et al.
Published: (2025)
by: Li, Zherui, et al.
Published: (2025)
Anchor-based Large Language Models
by: Pang, Jianhui, et al.
Published: (2024)
by: Pang, Jianhui, et al.
Published: (2024)
Learning from "Silly" Questions Improves Large Language Models, But Only Slightly
by: Zhu, Tingyuan, et al.
Published: (2024)
by: Zhu, Tingyuan, et al.
Published: (2024)
Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
by: Pang, Jianhui, et al.
Published: (2024)
by: Pang, Jianhui, et al.
Published: (2024)
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Similar Items
-
Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
by: Huang, Yuyi, et al.
Published: (2025) -
Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model
by: Xu, Haoyun, et al.
Published: (2024) -
Rethinking Prompt-based Debiasing in Large Language Models
by: Yang, Xinyi, et al.
Published: (2025) -
Prefix Text as a Yarn: Eliciting Non-English Alignment in Foundation Language Model
by: Zhan, Runzhe, et al.
Published: (2024) -
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
by: Zhan, Runzhe, et al.
Published: (2025)