Soft-Label Integration for Robust Toxicity Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Zelei, Wu, Xian, Yu, Jiahao, Han, Shuo, Cai, Xin-Qiang, Xing, Xinyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation
by: Cheng, Zelei, et al.
Published: (2024)
by: Cheng, Zelei, et al.
Published: (2024)
BlockScan: Detecting Anomalies in Blockchain Transactions
by: Yu, Jiahao, et al.
Published: (2024)
by: Yu, Jiahao, et al.
Published: (2024)
RMSL: Weakly-Supervised Insider Threat Detection with Robust Multi-sphere Learning
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
by: Kirci, Onur Alp, et al.
Published: (2025)
by: Kirci, Onur Alp, et al.
Published: (2025)
Training on Fake Labels: Mitigating Label Leakage in Split Learning via Secure Dimension Transformation
by: Jiang, Yukun, et al.
Published: (2024)
by: Jiang, Yukun, et al.
Published: (2024)
LAMD: Context-driven Android Malware Detection and Classification with LLMs
by: Qian, Xingzhi, et al.
Published: (2025)
by: Qian, Xingzhi, et al.
Published: (2025)
Towards Building a Robust Toxicity Predictor
by: Bespalov, Dmitriy, et al.
Published: (2024)
by: Bespalov, Dmitriy, et al.
Published: (2024)
Explainable Transformer-Based Email Phishing Classification with Adversarial Robustness
by: P, Sajad U
Published: (2025)
by: P, Sajad U
Published: (2025)
Free Lunch for Federated Remote Sensing Target Fine-Grained Classification: A Parameter-Efficient Framework
by: Chen, Shengchao, et al.
Published: (2024)
by: Chen, Shengchao, et al.
Published: (2024)
Advanced Payment Security System:XGBoost, LightGBM and SMOTE Integrated
by: Zheng, Qi, et al.
Published: (2024)
by: Zheng, Qi, et al.
Published: (2024)
VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
by: Yu, Simon, et al.
Published: (2025)
by: Yu, Simon, et al.
Published: (2025)
Beyond Detection: A Comprehensive Benchmark and Study on Representation Learning for Fine-Grained Webshell Family Classification
by: Han, Feijiang
Published: (2025)
by: Han, Feijiang
Published: (2025)
SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs
by: Pan, Chao, et al.
Published: (2026)
by: Pan, Chao, et al.
Published: (2026)
A Survey on Explainable Deep Reinforcement Learning
by: Cheng, Zelei, et al.
Published: (2025)
by: Cheng, Zelei, et al.
Published: (2025)
FedTracker: Furnishing Ownership Verification and Traceability for Federated Learning Model
by: Shao, Shuo, et al.
Published: (2022)
by: Shao, Shuo, et al.
Published: (2022)
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
by: Li, Xiaochen, et al.
Published: (2024)
by: Li, Xiaochen, et al.
Published: (2024)
Mixture of Robust Experts (MoRE):A Robust Denoising Method towards multiple perturbations
by: Cheng, Hao, et al.
Published: (2021)
by: Cheng, Hao, et al.
Published: (2021)
Exploring the Robustness of In-Context Learning with Noisy Labels
by: Cheng, Chen, et al.
Published: (2024)
by: Cheng, Chen, et al.
Published: (2024)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models
by: Luo, Haozheng, et al.
Published: (2025)
by: Luo, Haozheng, et al.
Published: (2025)
Local Data Quantity-Aware Weighted Averaging for Federated Learning with Dishonest Clients
by: Wu, Leming, et al.
Published: (2025)
by: Wu, Leming, et al.
Published: (2025)
Toxicity Detection towards Adaptability to Changing Perturbations
by: Kang, Hankun, et al.
Published: (2024)
by: Kang, Hankun, et al.
Published: (2024)
Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models
by: Zhang, Fuyao, et al.
Published: (2025)
by: Zhang, Fuyao, et al.
Published: (2025)
Differentially Private Worst-group Risk Minimization
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
OCGEC: One-class Graph Embedding Classification for DNN Backdoor Detection
by: Jiang, Haoyu, et al.
Published: (2023)
by: Jiang, Haoyu, et al.
Published: (2023)
SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
by: Yang, Wenyuan, et al.
Published: (2025)
by: Yang, Wenyuan, et al.
Published: (2025)
A Graph Transformer-Driven Approach for Network Robustness Learning
by: Zhang, Yu, et al.
Published: (2023)
by: Zhang, Yu, et al.
Published: (2023)
Are Robust LLM Fingerprints Adversarially Robust?
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
by: Li, Fengpeng, et al.
Published: (2026)
by: Li, Fengpeng, et al.
Published: (2026)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
by: Chen, Depeng, et al.
Published: (2024)
by: Chen, Depeng, et al.
Published: (2024)
Traversal-as-Policy: Log-Distilled Gated Behavior Trees as Externalized, Verifiable Policies for Safe, Robust, and Efficient Agents
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
Backdoor Graph Condensation
by: Wu, Jiahao, et al.
Published: (2024)
by: Wu, Jiahao, et al.
Published: (2024)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
by: Duan, Kaiwen, et al.
Published: (2025)
by: Duan, Kaiwen, et al.
Published: (2025)
Probing Network Decisions: Capturing Uncertainties and Unveiling Vulnerabilities Without Label Information
by: Joung, Youngju, et al.
Published: (2025)
by: Joung, Youngju, et al.
Published: (2025)
Similarity-based Label Inference Attack against Training and Inference of Split Learning
by: Liu, Junlin, et al.
Published: (2022)
by: Liu, Junlin, et al.
Published: (2022)
PrivTune: Efficient and Privacy-Preserving Fine-Tuning of Large Language Models via Device-Cloud Collaboration
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
FedCAP: Robust Federated Learning via Customized Aggregation and Personalization
by: Li, Youpeng, et al.
Published: (2024)
by: Li, Youpeng, et al.
Published: (2024)
Revolutionizing Encrypted Traffic Classification with MH-Net: A Multi-View Heterogeneous Graph Model
by: Zhang, Haozhen, et al.
Published: (2025)
by: Zhang, Haozhen, et al.
Published: (2025)
Robust Privacy: Inference-Time Privacy through Certified Robustness
by: Jin, Jiankai, et al.
Published: (2026)
by: Jin, Jiankai, et al.
Published: (2026)
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
by: Lu, Ning, et al.
Published: (2025)
by: Lu, Ning, et al.
Published: (2025)
Similar Items
-
RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation
by: Cheng, Zelei, et al.
Published: (2024) -
BlockScan: Detecting Anomalies in Blockchain Transactions
by: Yu, Jiahao, et al.
Published: (2024) -
RMSL: Weakly-Supervised Insider Threat Detection with Robust Multi-sphere Learning
by: Wang, Yang, et al.
Published: (2025) -
Strategic Sample Selection for Improved Clean-Label Backdoor Attacks in Text Classification
by: Kirci, Onur Alp, et al.
Published: (2025) -
Training on Fake Labels: Mitigating Label Leakage in Split Learning via Secure Dimension Transformation
by: Jiang, Yukun, et al.
Published: (2024)