ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Boyang, Shou, Hongzhe, Liang, Yuanyuan, Zhang, Jingbin, Zhou, Fang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STATE ToxiCN: A Benchmark for Span-level Target-Aware Toxicity Extraction in Chinese Hate Speech Detection
by: Bai, Zewen, et al.
Published: (2025)
by: Bai, Zewen, et al.
Published: (2025)
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
by: Xiao, Yunze, et al.
Published: (2024)
by: Xiao, Yunze, et al.
Published: (2024)
ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
by: Hui, Zheng, et al.
Published: (2024)
by: Hui, Zheng, et al.
Published: (2024)
ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection
by: Delaval, Axel, et al.
Published: (2025)
by: Delaval, Axel, et al.
Published: (2025)
A Survey of Explainable Knowledge Tracing
by: Bai, Yanhong, et al.
Published: (2024)
by: Bai, Yanhong, et al.
Published: (2024)
ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information
by: Hui, Zheng, et al.
Published: (2024)
by: Hui, Zheng, et al.
Published: (2024)
Reproducibility Report: Test-Time Training on Nearest Neighbors for Large Language Models
by: Zhou, Boyang, et al.
Published: (2025)
by: Zhou, Boyang, et al.
Published: (2025)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
by: Yang, Shujian, et al.
Published: (2025)
by: Yang, Shujian, et al.
Published: (2025)
Explainable Few-shot Knowledge Tracing
by: Li, Haoxuan, et al.
Published: (2024)
by: Li, Haoxuan, et al.
Published: (2024)
ToxiTwitch: Toward Emote-Aware Hybrid Moderation for Live Streaming Platforms
by: Ansari, Baktash, et al.
Published: (2026)
by: Ansari, Baktash, et al.
Published: (2026)
Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites
by: Wang, Xintong, et al.
Published: (2025)
by: Wang, Xintong, et al.
Published: (2025)
Aligned Probing: Relating Toxic Behavior and Model Internals
by: Waldis, Andreas, et al.
Published: (2025)
by: Waldis, Andreas, et al.
Published: (2025)
Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon
by: Ma, Xuchen, et al.
Published: (2025)
by: Ma, Xuchen, et al.
Published: (2025)
Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks
by: Bai, Zewen, et al.
Published: (2025)
by: Bai, Zewen, et al.
Published: (2025)
HRDE: Retrieval-Augmented Large Language Models for Chinese Health Rumor Detection and Explainability
by: Chen, Yanfang, et al.
Published: (2024)
by: Chen, Yanfang, et al.
Published: (2024)
LLM-KT: Aligning Large Language Models with Knowledge Tracing using a Plug-and-Play Instruction
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
Towards Training A Chinese Large Language Model for Anesthesiology
by: Wang, Zhonghai, et al.
Published: (2024)
by: Wang, Zhonghai, et al.
Published: (2024)
Chinese-SkillSpan: A Span-Level Dataset for ESCO-Aligned Competency Extraction from Chinese Job Ads
by: Li, Guojing, et al.
Published: (2026)
by: Li, Guojing, et al.
Published: (2026)
CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark
by: Zhang, Junzhao, et al.
Published: (2026)
by: Zhang, Junzhao, et al.
Published: (2026)
A Training-free LLM-based Approach to General Chinese Character Error Correction
by: Zhou, Houquan, et al.
Published: (2025)
by: Zhou, Houquan, et al.
Published: (2025)
Toxicity Detection for Free
by: Hu, Zhanhao, et al.
Published: (2024)
by: Hu, Zhanhao, et al.
Published: (2024)
Toxicity Detection towards Adaptability to Changing Perturbations
by: Kang, Hankun, et al.
Published: (2024)
by: Kang, Hankun, et al.
Published: (2024)
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
by: Wang, Shaojie, et al.
Published: (2026)
by: Wang, Shaojie, et al.
Published: (2026)
Aligning Teacher with Student Preferences for Tailored Training Data Generation
by: Liu, Yantao, et al.
Published: (2024)
by: Liu, Yantao, et al.
Published: (2024)
Prior Constraints-based Reward Model Training for Aligning Large Language Models
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
Explainable Depression Detection in Clinical Interviews with Personalized Retrieval-Augmented Generation
by: Zhang, Linhai, et al.
Published: (2025)
by: Zhang, Linhai, et al.
Published: (2025)
Speculating LLMs' Chinese Training Data Pollution from Their Tokens
by: Zhang, Qingjie, et al.
Published: (2025)
by: Zhang, Qingjie, et al.
Published: (2025)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
by: Zhao, Yibo, et al.
Published: (2024)
by: Zhao, Yibo, et al.
Published: (2024)
From Noisy Traces to Stable Gradients: Bias-Variance Optimized Preference Optimization for Aligning Large Reasoning Models
by: Zhu, Mingkang, et al.
Published: (2025)
by: Zhu, Mingkang, et al.
Published: (2025)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
Toxic HallucinAItions: Perturbing Prompts and Tracing LLM Circuits
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
RAIR: Retrieval-Augmented Iterative Refinement for Chinese Spelling Correction
by: Liang, Junhong, et al.
Published: (2025)
by: Liang, Junhong, et al.
Published: (2025)
A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language Models
by: Zhou, Houquan, et al.
Published: (2024)
by: Zhou, Houquan, et al.
Published: (2024)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Sparse Attention across Multiple-context KV Cache
by: Cao, Ziyi, et al.
Published: (2025)
by: Cao, Ziyi, et al.
Published: (2025)
EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error Correction
by: Ye, Jingheng, et al.
Published: (2024)
by: Ye, Jingheng, et al.
Published: (2024)
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting
by: Kang, Jingyi, et al.
Published: (2026)
by: Kang, Jingyi, et al.
Published: (2026)
Similar Items
-
STATE ToxiCN: A Benchmark for Span-level Target-Aware Toxicity Extraction in Chinese Hate Speech Detection
by: Bai, Zewen, et al.
Published: (2025) -
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
by: Xiao, Yunze, et al.
Published: (2024) -
ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation
by: Li, Peiran, et al.
Published: (2026) -
ToxiLab: How Well Do Open-Source LLMs Generate Synthetic Toxicity Data?
by: Hui, Zheng, et al.
Published: (2024) -
ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection
by: Delaval, Axel, et al.
Published: (2025)