GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Youliang, Jiao, Wenxiang, Wang, Wenxuan, Huang, Jen-tse, He, Pinjia, Shi, Shuming, Tu, Zhaopeng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024)
by: Yuan, Youliang, et al.
Published: (2024)
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench
by: Huang, Jen-tse, et al.
Published: (2023)
by: Huang, Jen-tse, et al.
Published: (2023)
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
by: Liu, Xiaoyuan, et al.
Published: (2024)
by: Liu, Xiaoyuan, et al.
Published: (2024)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
All Languages Matter: On the Multilingual Safety of Large Language Models
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
by: Yuan, Youliang, et al.
Published: (2025)
by: Yuan, Youliang, et al.
Published: (2025)
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
On the Shortcut Learning in Multilingual Neural Machine Translation
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
by: Huang, Jen-Tse, et al.
Published: (2025)
by: Huang, Jen-Tse, et al.
Published: (2025)
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
by: Huang, Jen-tse, et al.
Published: (2024)
by: Huang, Jen-tse, et al.
Published: (2024)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
by: Yuan, Youliang, et al.
Published: (2025)
by: Yuan, Youliang, et al.
Published: (2025)
Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
by: Huang, Jen-tse, et al.
Published: (2023)
by: Huang, Jen-tse, et al.
Published: (2023)
Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model
by: He, Zhiwei, et al.
Published: (2024)
by: He, Zhiwei, et al.
Published: (2024)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
by: Zhao, Sihang, et al.
Published: (2024)
by: Zhao, Sihang, et al.
Published: (2024)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
Learning to Ask: When LLM Agents Meet Unclear Instruction
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
by: Zhao, Sihang, et al.
Published: (2026)
by: Zhao, Sihang, et al.
Published: (2026)
Revisiting the Reliability of Psychological Scales on Large Language Models
by: Huang, Jen-tse, et al.
Published: (2023)
by: Huang, Jen-tse, et al.
Published: (2023)
Understanding and Mitigating the Uncertainty in Zero-Shot Translation
by: Wang, Wenxuan, et al.
Published: (2022)
by: Wang, Wenxuan, et al.
Published: (2022)
Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
by: Liang, Tian, et al.
Published: (2023)
by: Liang, Tian, et al.
Published: (2023)
Benchmarking LLMs via Uncertainty Quantification
by: Ye, Fanghua, et al.
Published: (2024)
by: Ye, Fanghua, et al.
Published: (2024)
NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates
by: Deng, Hexuan, et al.
Published: (2024)
by: Deng, Hexuan, et al.
Published: (2024)
Cross-modality Data Augmentation for End-to-End Sign Language Translation
by: Ye, Jinhui, et al.
Published: (2023)
by: Ye, Jinhui, et al.
Published: (2023)
On the Failure of Latent State Persistence in Large Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
by: Liu, Xiaoyuan, et al.
Published: (2025)
by: Liu, Xiaoyuan, et al.
Published: (2025)
New Job, New Gender? Measuring the Social Bias in Image Generation Models
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
by: Du, Yongkang, et al.
Published: (2025)
by: Du, Yongkang, et al.
Published: (2025)
DRPruning: Efficient Large Language Model Pruning through Distributionally Robust Optimization
by: Deng, Hexuan, et al.
Published: (2024)
by: Deng, Hexuan, et al.
Published: (2024)
How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO
by: Ng, Man Tik, et al.
Published: (2024)
by: Ng, Man Tik, et al.
Published: (2024)
CoAct: A Global-Local Hierarchy for Autonomous Agent Collaboration
by: Hou, Xinming, et al.
Published: (2024)
by: Hou, Xinming, et al.
Published: (2024)
AuditGPT: Auditing Smart Contracts with ChatGPT
by: Xia, Shihao, et al.
Published: (2024)
by: Xia, Shihao, et al.
Published: (2024)
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
by: Yi, Zihao, et al.
Published: (2025)
by: Yi, Zihao, et al.
Published: (2025)
GroupGPT: A Token-efficient and Privacy-preserving Agentic Framework for Multi-User Chat Assistant
by: Shen, Zhuokang, et al.
Published: (2026)
by: Shen, Zhuokang, et al.
Published: (2026)
AI Sees Your Location, But With A Bias Toward The Wealthy World
by: Huang, Jingyuan, et al.
Published: (2025)
by: Huang, Jingyuan, et al.
Published: (2025)
Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation
by: Huang, Jen-tse, et al.
Published: (2026)
by: Huang, Jen-tse, et al.
Published: (2026)
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
by: Huang, Jen-tse, et al.
Published: (2026)
by: Huang, Jen-tse, et al.
Published: (2026)
Unsupervised Sign Language Translation and Generation
by: Guo, Zhengsheng, et al.
Published: (2024)
by: Guo, Zhengsheng, et al.
Published: (2024)
Similar Items
-
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024) -
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench
by: Huang, Jen-tse, et al.
Published: (2023) -
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
by: Liu, Xiaoyuan, et al.
Published: (2024) -
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025) -
All Languages Matter: On the Multilingual Safety of Large Language Models
by: Wang, Wenxuan, et al.
Published: (2023)