Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
Fuente:
arXiv
Guardado en:
| Autores principales: | Stepanov, Ihor, Smechov, Aleksandr |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GLiNER multi-task: Generalist Lightweight Model for Various Information Extraction Tasks
por: Stepanov, Ihor, et al.
Publicado: (2024)
por: Stepanov, Ihor, et al.
Publicado: (2024)
GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
por: Stepanov, Ihor, et al.
Publicado: (2025)
por: Stepanov, Ihor, et al.
Publicado: (2025)
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
por: Almohaimeed, Saad, et al.
Publicado: (2025)
por: Almohaimeed, Saad, et al.
Publicado: (2025)
Efficient Safety Retrofitting Against Jailbreaking for LLMs
por: Garcia-Gasulla, Dario, et al.
Publicado: (2025)
por: Garcia-Gasulla, Dario, et al.
Publicado: (2025)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
por: Berezin, Sergei, et al.
Publicado: (2025)
por: Berezin, Sergei, et al.
Publicado: (2025)
Improving Hate Speech Classification with Cross-Taxonomy Dataset Integration
por: Fillies, Jan, et al.
Publicado: (2025)
por: Fillies, Jan, et al.
Publicado: (2025)
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI
por: Böck, Adrian Jaques, et al.
Publicado: (2024)
por: Böck, Adrian Jaques, et al.
Publicado: (2024)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
por: Chan, Yik Siu, et al.
Publicado: (2025)
por: Chan, Yik Siu, et al.
Publicado: (2025)
Dealing with Annotator Disagreement in Hate Speech Classification
por: Dehghan, Somaiyeh, et al.
Publicado: (2025)
por: Dehghan, Somaiyeh, et al.
Publicado: (2025)
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
por: Das, Amit, et al.
Publicado: (2024)
por: Das, Amit, et al.
Publicado: (2024)
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
por: Zhao, Chengshuai, et al.
Publicado: (2025)
por: Zhao, Chengshuai, et al.
Publicado: (2025)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
por: AlDahoul, Nouar, et al.
Publicado: (2025)
por: AlDahoul, Nouar, et al.
Publicado: (2025)
Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
por: Islam, Akif, et al.
Publicado: (2025)
por: Islam, Akif, et al.
Publicado: (2025)
Transformers and Ensemble methods: A solution for Hate Speech Detection in Arabic languages
por: de Paula, Angel Felipe Magnossão, et al.
Publicado: (2023)
por: de Paula, Angel Felipe Magnossão, et al.
Publicado: (2023)
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
por: Nirmal, Ayushi, et al.
Publicado: (2024)
por: Nirmal, Ayushi, et al.
Publicado: (2024)
Large Language Models in the Task of Automatic Validation of Text Classifier Predictions
por: Tsymbalov, Aleksandr, et al.
Publicado: (2025)
por: Tsymbalov, Aleksandr, et al.
Publicado: (2025)
Identifying False Content and Hate Speech in Sinhala YouTube Videos by Analyzing the Audio
por: Wickramaarachchi, W. A. K. M., et al.
Publicado: (2024)
por: Wickramaarachchi, W. A. K. M., et al.
Publicado: (2024)
PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
por: Rao, Akshaj Prashanth, et al.
Publicado: (2025)
por: Rao, Akshaj Prashanth, et al.
Publicado: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
por: Chaudhary, Maheep, et al.
Publicado: (2025)
por: Chaudhary, Maheep, et al.
Publicado: (2025)
Multilingual Hate Speech Detection in Social Media Using Translation-Based Approaches with Large Language Models
por: Usman, Muhammad, et al.
Publicado: (2025)
por: Usman, Muhammad, et al.
Publicado: (2025)
1-800-SHARED-TASKS @ NLU of Devanagari Script Languages: Detection of Language, Hate Speech, and Targets using LLMs
por: Purbey, Jebish, et al.
Publicado: (2024)
por: Purbey, Jebish, et al.
Publicado: (2024)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
por: Oh, Sejoon, et al.
Publicado: (2024)
por: Oh, Sejoon, et al.
Publicado: (2024)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
por: Cheng, Yixin, et al.
Publicado: (2024)
por: Cheng, Yixin, et al.
Publicado: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
NoisyHate: Mining Online Human-Written Perturbations for Realistic Robustness Benchmarking of Content Moderation Models
por: Ye, Yiran, et al.
Publicado: (2023)
por: Ye, Yiran, et al.
Publicado: (2023)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
por: Pandey, Punya Syon, et al.
Publicado: (2025)
por: Pandey, Punya Syon, et al.
Publicado: (2025)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
por: Liu, Kangwei, et al.
Publicado: (2025)
por: Liu, Kangwei, et al.
Publicado: (2025)
Jailbreaking with Universal Multi-Prompts
por: Hsu, Yu-Ling, et al.
Publicado: (2025)
por: Hsu, Yu-Ling, et al.
Publicado: (2025)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
por: Gringras, David
Publicado: (2026)
por: Gringras, David
Publicado: (2026)
Best-of-N Jailbreaking
por: Hughes, John, et al.
Publicado: (2024)
por: Hughes, John, et al.
Publicado: (2024)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
por: Lee, Isack, et al.
Publicado: (2024)
por: Lee, Isack, et al.
Publicado: (2024)
Whitening Not Recommended for Classification Tasks in LLMs
por: Forooghi, Ali, et al.
Publicado: (2024)
por: Forooghi, Ali, et al.
Publicado: (2024)
OSPC: Artificial VLM Features for Hateful Meme Detection
por: Grönquist, Peter
Publicado: (2024)
por: Grönquist, Peter
Publicado: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
por: Chen, Taiye, et al.
Publicado: (2025)
por: Chen, Taiye, et al.
Publicado: (2025)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
por: Dinuta, Eduard Stefan, et al.
Publicado: (2025)
por: Dinuta, Eduard Stefan, et al.
Publicado: (2025)
Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
por: Neill, James O', et al.
Publicado: (2025)
por: Neill, James O', et al.
Publicado: (2025)
OrchMoE: Efficient Multi-Adapter Learning with Task-Skill Synergy
por: Wang, Haowen, et al.
Publicado: (2024)
por: Wang, Haowen, et al.
Publicado: (2024)
Lightweight Safety Classification Using Pruned Language Models
por: Sawtell, Mason, et al.
Publicado: (2024)
por: Sawtell, Mason, et al.
Publicado: (2024)
Keyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content
por: Rizvi, F. A., et al.
Publicado: (2025)
por: Rizvi, F. A., et al.
Publicado: (2025)
Mitigating Extrinsic Gender Bias for Bangla Classification Tasks
por: Joy, Sajib Kumar Saha, et al.
Publicado: (2024)
por: Joy, Sajib Kumar Saha, et al.
Publicado: (2024)
Ejemplares similares
-
GLiNER multi-task: Generalist Lightweight Model for Various Information Extraction Tasks
por: Stepanov, Ihor, et al.
Publicado: (2024) -
GLiClass: Generalist Lightweight Model for Sequence Classification Tasks
por: Stepanov, Ihor, et al.
Publicado: (2025) -
Towards Generalizable Generic Harmful Speech Datasets for Implicit Hate Speech Detection
por: Almohaimeed, Saad, et al.
Publicado: (2025) -
Efficient Safety Retrofitting Against Jailbreaking for LLMs
por: Garcia-Gasulla, Dario, et al.
Publicado: (2025) -
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
por: Berezin, Sergei, et al.
Publicado: (2025)