ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Yunze, Hu, Yujia, Choo, Kenny Tsu Wei, Lee, Roy Ka-wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lost in Pronunciation: Detecting Chinese Offensive Language Disguised by Phonetic Cloaking Replacement
by: Guo, Haotan, et al.
Published: (2025)
by: Guo, Haotan, et al.
Published: (2025)
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
by: Ma, Xuchen, et al.
Published: (2025)
by: Ma, Xuchen, et al.
Published: (2025)
MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations
by: Xue, Qiyao, et al.
Published: (2025)
by: Xue, Qiyao, et al.
Published: (2025)
Chinese Offensive Language Detection:Current Status and Future Directions
by: Xiao, Yunze, et al.
Published: (2024)
by: Xiao, Yunze, et al.
Published: (2024)
SGHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Singapore
by: Ng, Ri Chi, et al.
Published: (2024)
by: Ng, Ri Chi, et al.
Published: (2024)
Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
by: Ryan, Yuriel, et al.
Published: (2025)
by: Ryan, Yuriel, et al.
Published: (2025)
"Is This Really a Human Peer Supporter?": Misalignments Between Peer Supporters and Experts in LLM-Supported Interactions
by: Sim, Kellie Yu Hui, et al.
Published: (2025)
by: Sim, Kellie Yu Hui, et al.
Published: (2025)
JiraiBench: A Bilingual Benchmark for Evaluating Large Language Models' Detection of Human Self-Destructive Behavior Content in Jirai Community
by: Xiao, Yunze, et al.
Published: (2025)
by: Xiao, Yunze, et al.
Published: (2025)
3D Invisible Cloak
by: Xue, Mingfu, et al.
Published: (2020)
by: Xue, Mingfu, et al.
Published: (2020)
Designing for Novice Debuggers: A Pilot Study on an AI-Assisted Debugging Tool
by: Kurniawan, Oka, et al.
Published: (2025)
by: Kurniawan, Oka, et al.
Published: (2025)
ID-Cloak: Crafting Identity-Specific Cloaks Against Personalized Text-to-Image Generation
by: Teng, Qianrui, et al.
Published: (2025)
by: Teng, Qianrui, et al.
Published: (2025)
Unmasking Implicit Bias: Evaluating Persona-Prompted LLM Responses in Power-Disparate Social Scenarios
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
A Taxonomy of Human--MLLM Interaction in Early-Stage Sketch-Based Design Ideation
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
More Than 1v1: Human-AI Alignment in Early Developmental Communities with Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
by: Alipour, Shayan, et al.
Published: (2024)
by: Alipour, Shayan, et al.
Published: (2024)
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Detecting Offensive Cyber Agents: A Detection-in-Depth Approach
by: Mittelsteadt, Matt, et al.
Published: (2026)
by: Mittelsteadt, Matt, et al.
Published: (2026)
Structured Semantic Cloaking for Jailbreak Attacks on Large Language Models
by: Sun, Xiaobing, et al.
Published: (2026)
by: Sun, Xiaobing, et al.
Published: (2026)
HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations
by: Hu, Yujia, et al.
Published: (2026)
by: Hu, Yujia, et al.
Published: (2026)
Geometry Cloak: Preventing TGS-based 3D Reconstruction from Copyrighted Images
by: Song, Qi, et al.
Published: (2024)
by: Song, Qi, et al.
Published: (2024)
ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection
by: Delaval, Axel, et al.
Published: (2025)
by: Delaval, Axel, et al.
Published: (2025)
FaceCloak: Learning to Protect Face Templates
by: Banerjee, Sudipta, et al.
Published: (2025)
by: Banerjee, Sudipta, et al.
Published: (2025)
STATE ToxiCN: A Benchmark for Span-level Target-Aware Toxicity Extraction in Chinese Hate Speech Detection
by: Bai, Zewen, et al.
Published: (2025)
by: Bai, Zewen, et al.
Published: (2025)
Cloaked Classifiers: Pseudonymization Strategies on Sensitive Classification Tasks
by: Riabi, Arij, et al.
Published: (2024)
by: Riabi, Arij, et al.
Published: (2024)
Examining Augmented Virtuality Impairment Simulation for Mobile App Accessibility Design
by: Choo, Kenny Tsu Wei, et al.
Published: (2025)
by: Choo, Kenny Tsu Wei, et al.
Published: (2025)
Towards Multimodal Large-Language Models for Parent-Child Interaction: A Focus on Joint Attention
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
NullSwap: Proactive Identity Cloaking Against Deepfake Face Swapping
by: Wang, Tianyi, et al.
Published: (2025)
by: Wang, Tianyi, et al.
Published: (2025)
Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
Understanding Computer Science Students' Career Fair Experiences: Goals, Preparation, and Outcomes
by: Lee, Briana, et al.
Published: (2025)
by: Lee, Briana, et al.
Published: (2025)
Foreign Domestic Workers' Perspectives on an LLM-Based Emotional Support tool for Caregiving Burden
by: Teng, Shin Shoon Nicholas, et al.
Published: (2026)
by: Teng, Shin Shoon Nicholas, et al.
Published: (2026)
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
by: Morabito, Robert, et al.
Published: (2024)
by: Morabito, Robert, et al.
Published: (2024)
Embracing Contradiction: Theoretical Inconsistency Will Not Impede the Road of Building Responsible AI Systems
by: Dai, Gordon, et al.
Published: (2025)
by: Dai, Gordon, et al.
Published: (2025)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
by: Lu, Junyu, et al.
Published: (2025)
by: Lu, Junyu, et al.
Published: (2025)
AustroTox: A Dataset for Target-Based Austrian German Offensive Language Detection
by: Pachinger, Pia, et al.
Published: (2024)
by: Pachinger, Pia, et al.
Published: (2024)
FairPair: A Robust Evaluation of Biases in Language Models through Paired Perturbations
by: Dwivedi-Yu, Jane, et al.
Published: (2024)
by: Dwivedi-Yu, Jane, et al.
Published: (2024)
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
MetaCloak-JPEG: JPEG-Robust Adversarial Perturbation for Preventing Unauthorized DreamBooth-Based Deepfake Generation
by: Fardin, Tanjim Rahaman, et al.
Published: (2026)
by: Fardin, Tanjim Rahaman, et al.
Published: (2026)
Similar Items
-
Lost in Pronunciation: Detecting Chinese Offensive Language Disguised by Phonetic Cloaking Replacement
by: Guo, Haotan, et al.
Published: (2025) -
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025) -
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
by: Ma, Xuchen, et al.
Published: (2025) -
MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations
by: Xue, Qiyao, et al.
Published: (2025) -
Chinese Offensive Language Detection:Current Status and Future Directions
by: Xiao, Yunze, et al.
Published: (2024)