When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Cui, Shiyao, Feng, Xijia, Wang, Yingkang, Yang, Junxiao, Zhang, Zhexin, Sikdar, Biplab, Wang, Hongning, Qiu, Han, Huang, Minlie |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
par: Yang, Junxiao, et autres
Publié: (2025)
par: Yang, Junxiao, et autres
Publié: (2025)
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
par: Zhang, Zhexin, et autres
Publié: (2025)
par: Zhang, Zhexin, et autres
Publié: (2025)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
par: Zhang, Zhexin, et autres
Publié: (2024)
par: Zhang, Zhexin, et autres
Publié: (2024)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
par: Cui, Shiyao, et autres
Publié: (2025)
par: Cui, Shiyao, et autres
Publié: (2025)
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
par: Zhang, Zhexin, et autres
Publié: (2025)
par: Zhang, Zhexin, et autres
Publié: (2025)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
par: Zhang, Zhexin, et autres
Publié: (2024)
par: Zhang, Zhexin, et autres
Publié: (2024)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
par: Zhang, Zhexin, et autres
Publié: (2023)
par: Zhang, Zhexin, et autres
Publié: (2023)
The Superalignment of Superhuman Intelligence with Large Language Models
par: Huang, Minlie, et autres
Publié: (2024)
par: Huang, Minlie, et autres
Publié: (2024)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
par: Yang, Junxiao, et autres
Publié: (2026)
par: Yang, Junxiao, et autres
Publié: (2026)
The Missing Half: Unveiling Training-time Implicit Safety Risks Beyond Deployment
par: Zhang, Zhexin, et autres
Publié: (2026)
par: Zhang, Zhexin, et autres
Publié: (2026)
BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs
par: Yang, Junxiao, et autres
Publié: (2025)
par: Yang, Junxiao, et autres
Publié: (2025)
LongSafety: Evaluating Long-Context Safety of Large Language Models
par: Lu, Yida, et autres
Publié: (2025)
par: Lu, Yida, et autres
Publié: (2025)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
par: Yang, Shujian, et autres
Publié: (2025)
par: Yang, Shujian, et autres
Publié: (2025)
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
par: Chen, Renmiao, et autres
Publié: (2025)
par: Chen, Renmiao, et autres
Publié: (2025)
SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
par: Tu, Jinzhe, et autres
Publié: (2026)
par: Tu, Jinzhe, et autres
Publié: (2026)
The Effects of Emoji in Social Media Food Content: How Smileys, Hearts, and Hamburgers Affect Engagement, Attitudes, and Intentions
par: Lieke Verheijen, et autres
Publié: (2025)
par: Lieke Verheijen, et autres
Publié: (2025)
AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement
par: Zhang, Zhexin, et autres
Publié: (2025)
par: Zhang, Zhexin, et autres
Publié: (2025)
Language Model Decoding as Direct Metrics Optimization
par: Ji, Haozhe, et autres
Publié: (2023)
par: Ji, Haozhe, et autres
Publié: (2023)
Grounding LLMs in Scientific Discovery via Embodied Actions
par: Zhang, Bo, et autres
Publié: (2026)
par: Zhang, Bo, et autres
Publié: (2026)
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
par: Zhang, Zhexin, et autres
Publié: (2024)
par: Zhang, Zhexin, et autres
Publié: (2024)
LiteAtt: A Peer-to-Peer Self-Attestation Framework and Handshake Protocol for Connected IoT Devices
par: Kohli, Varun, et autres
Publié: (2026)
par: Kohli, Varun, et autres
Publié: (2026)
PEANUT: Perturbations by Eigenvector Alignment for Attacking Graph Neural Networks Under Topology-Driven Message Passing
par: Kohli, Bhavya, et autres
Publié: (2026)
par: Kohli, Bhavya, et autres
Publié: (2026)
From Text to Emoji: How PEFT-Driven Personality Manipulation Unleashes the Emoji Potential in LLMs
par: Jain, Navya, et autres
Publié: (2024)
par: Jain, Navya, et autres
Publié: (2024)
Trust-Region Adaptive Policy Optimization
par: Su, Mingyu, et autres
Publié: (2025)
par: Su, Mingyu, et autres
Publié: (2025)
Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning
par: Wen, Jiaxin, et autres
Publié: (2024)
par: Wen, Jiaxin, et autres
Publié: (2024)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
par: Cui, Shiyao, et autres
Publié: (2023)
par: Cui, Shiyao, et autres
Publié: (2023)
When In‐Groups Turn Hostile: Criminal Record, Intergroup Bias, and the Black Sheep Effect in Eyewitness Judgments
par: Nir Rozmann
Publié: (2026)
par: Nir Rozmann
Publié: (2026)
RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
par: Feng, Andrew Zhuoer, et autres
Publié: (2026)
par: Feng, Andrew Zhuoer, et autres
Publié: (2026)
Privacy-Preserving Collaborative Split Learning Framework for Smart Grid Load Forecasting
par: Iqbal, Asif, et autres
Publié: (2024)
par: Iqbal, Asif, et autres
Publié: (2024)
Learning Task Decomposition to Assist Humans in Competitive Programming
par: Wen, Jiaxin, et autres
Publié: (2024)
par: Wen, Jiaxin, et autres
Publié: (2024)
AMOR: A Recipe for Building Adaptable Modular Knowledge Agents Through Process Feedback
par: Guan, Jian, et autres
Publié: (2024)
par: Guan, Jian, et autres
Publié: (2024)
In-Memory Sorting-Searching with Cayley Tree
par: Paul, Subrata, et autres
Publié: (2025)
par: Paul, Subrata, et autres
Publié: (2025)
When the Air Turns Toxic, So Might Our Behavior
par: Vạc Hoa
Publié: (2025)
par: Vạc Hoa
Publié: (2025)
Think Socially via Cognitive Reasoning
par: Zhou, Jinfeng, et autres
Publié: (2025)
par: Zhou, Jinfeng, et autres
Publié: (2025)
UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion Models
par: Han, Yuning, et autres
Publié: (2024)
par: Han, Yuning, et autres
Publié: (2024)
The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
par: Chen, Renmiao, et autres
Publié: (2026)
par: Chen, Renmiao, et autres
Publié: (2026)
Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework
par: Zhang, Xuanming, et autres
Publié: (2024)
par: Zhang, Xuanming, et autres
Publié: (2024)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
par: Lu, Yiyang, et autres
Publié: (2026)
par: Lu, Yiyang, et autres
Publié: (2026)
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
par: Dongre, Vardhan, et autres
Publié: (2026)
par: Dongre, Vardhan, et autres
Publié: (2026)
Unified Framework for Qualifying Security Boundary of PUFs Against Machine Learning Attacks
par: Fei, Hongming, et autres
Publié: (2026)
par: Fei, Hongming, et autres
Publié: (2026)
Documents similaires
-
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
par: Yang, Junxiao, et autres
Publié: (2025) -
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
par: Zhang, Zhexin, et autres
Publié: (2025) -
Agent-SafetyBench: Evaluating the Safety of LLM Agents
par: Zhang, Zhexin, et autres
Publié: (2024) -
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
par: Cui, Shiyao, et autres
Publié: (2025) -
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
par: Zhang, Zhexin, et autres
Publié: (2025)