ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Kangwei, Cheng, Siyuan, Tian, Bozhong, Liang, Xiaozhuan, Yin, Yuyang, Han, Meng, Zhang, Ningyu, Hooi, Bryan, Chen, Xi, Deng, Shumin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
by: Belkhiter, Yannis, et al.
Published: (2024)
by: Belkhiter, Yannis, et al.
Published: (2024)
Supporting Human Raters with the Detection of Harmful Content using Large Language Models
by: Thomas, Kurt, et al.
Published: (2024)
by: Thomas, Kurt, et al.
Published: (2024)
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
by: Liu, Yuexiao, et al.
Published: (2025)
by: Liu, Yuexiao, et al.
Published: (2025)
Automated Phishing Detection Using URLs and Webpages
by: Wang, Huilin, et al.
Published: (2024)
by: Wang, Huilin, et al.
Published: (2024)
RobustMask: Certified Robustness against Adversarial Neural Ranking Attack via Randomized Masking
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
by: Yang, Jirui, et al.
Published: (2025)
by: Yang, Jirui, et al.
Published: (2025)
Poisoning Decentralized Collaborative Recommender System and Its Countermeasures
by: Zheng, Ruiqi, et al.
Published: (2024)
by: Zheng, Ruiqi, et al.
Published: (2024)
PhishIntel: Toward Practical Deployment of Reference-Based Phishing Detection
by: Li, Yuexin, et al.
Published: (2024)
by: Li, Yuexin, et al.
Published: (2024)
Harnessing TI Feeds for Exploitation Detection
by: Patel, Kajal, et al.
Published: (2024)
by: Patel, Kajal, et al.
Published: (2024)
Can It Reach the Generator? Investigating the Survival of Prompt-Injection Attacks in Realistic RAG Settings
by: Yin, Yu, et al.
Published: (2026)
by: Yin, Yu, et al.
Published: (2026)
Phishing Email Detection Using Large Language Models
by: Hasan, Najmul, et al.
Published: (2025)
by: Hasan, Najmul, et al.
Published: (2025)
Detecting Cryptographically Relevant Software Packages with Collaborative LLMs
by: Hirsch, Eduard, et al.
Published: (2026)
by: Hirsch, Eduard, et al.
Published: (2026)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
by: Cao, Tri, et al.
Published: (2025)
by: Cao, Tri, et al.
Published: (2025)
BERTDetect: A Neural Topic Modelling Approach for Android Malware Detection
by: Ranaweera, Nishavi, et al.
Published: (2025)
by: Ranaweera, Nishavi, et al.
Published: (2025)
Poisoning Attacks and Defenses in Recommender Systems: A Survey
by: Wang, Zongwei, et al.
Published: (2024)
by: Wang, Zongwei, et al.
Published: (2024)
BiRD: A Bidirectional Ranking Defense Mechanism for Retrieval Augmented Generation
by: Gao, Chengcai, et al.
Published: (2026)
by: Gao, Chengcai, et al.
Published: (2026)
CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis
by: Eswaran, Anushri, et al.
Published: (2026)
by: Eswaran, Anushri, et al.
Published: (2026)
HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking
by: Wang, Zeng, et al.
Published: (2026)
by: Wang, Zeng, et al.
Published: (2026)
SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model
by: Liang, Xun, et al.
Published: (2025)
by: Liang, Xun, et al.
Published: (2025)
Personalized w-Event Privacy for Infinite Stream Estimation
by: Du, Leilei, et al.
Published: (2026)
by: Du, Leilei, et al.
Published: (2026)
Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models
by: Gong, Yuyang, et al.
Published: (2025)
by: Gong, Yuyang, et al.
Published: (2025)
Towards A Unified View of Answer Calibration for Multi-Step Reasoning
by: Deng, Shumin, et al.
Published: (2023)
by: Deng, Shumin, et al.
Published: (2023)
Benchmarking Poisoning Attacks against Retrieval-Augmented Generation
by: Zhang, Baolei, et al.
Published: (2025)
by: Zhang, Baolei, et al.
Published: (2025)
Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
Unbundle-Rewrite-Rebundle: Runtime Detection and Rewriting of Privacy-Harming Code in JavaScript Bundles
by: Ali, Mir Masood, et al.
Published: (2024)
by: Ali, Mir Masood, et al.
Published: (2024)
GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
Versatile and Fast Location-Based Private Information Retrieval with Fully Homomorphic Encryption over the Torus
by: Yoo, Joon Soo, et al.
Published: (2025)
by: Yoo, Joon Soo, et al.
Published: (2025)
Adversarial Hubness in Multi-Modal Retrieval
by: Zhang, Tingwei, et al.
Published: (2024)
by: Zhang, Tingwei, et al.
Published: (2024)
ProveRAG: Provenance-Driven Vulnerability Analysis with Automated Retrieval-Augmented LLMs
by: Fayyazi, Reza, et al.
Published: (2024)
by: Fayyazi, Reza, et al.
Published: (2024)
First Steps, Lasting Impact: Platform-Aware Forensics for the Next Generation of Analysts
by: Jain, Vinayak, et al.
Published: (2026)
by: Jain, Vinayak, et al.
Published: (2026)
Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies
by: Srinath, Mukund, et al.
Published: (2020)
by: Srinath, Mukund, et al.
Published: (2020)
Green-Red Watermarking for Recommender Systems
by: Zhou, Lei, et al.
Published: (2026)
by: Zhou, Lei, et al.
Published: (2026)
RemoteRAG: A Privacy-Preserving LLM Cloud RAG Service
by: Cheng, Yihang, et al.
Published: (2024)
by: Cheng, Yihang, et al.
Published: (2024)
SoK: Timeline based event reconstruction for digital forensics: Terminology, methodology, and current challenges
by: Breitinger, Frank, et al.
Published: (2025)
by: Breitinger, Frank, et al.
Published: (2025)
Cybersecurity Data Extraction from Common Crawl
by: Mahara, Ashim
Published: (2025)
by: Mahara, Ashim
Published: (2025)
DV-FSR: A Dual-View Target Attack Framework for Federated Sequential Recommendation
by: Qin, Qitao, et al.
Published: (2024)
by: Qin, Qitao, et al.
Published: (2024)
Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacks
by: Bagwe, Gaurav, et al.
Published: (2025)
by: Bagwe, Gaurav, et al.
Published: (2025)
Similar Items
-
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024) -
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
by: Jiang, Yukun, et al.
Published: (2026) -
HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment
by: Belkhiter, Yannis, et al.
Published: (2024) -
Supporting Human Raters with the Detection of Harmful Content using Large Language Models
by: Thomas, Kurt, et al.
Published: (2024) -
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
by: Liu, Yuexiao, et al.
Published: (2025)