Covering Cracks in Content Moderation: Delexicalized Distant Supervision for Illicit Drug Jargon Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Song, Minkyoo, Jang, Eugene, Kim, Jaehan, Shin, Seungwon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm
di: Kim, Jaehan, et al.
Pubblicazione: (2024)
di: Kim, Jaehan, et al.
Pubblicazione: (2024)
Claim-Guided Textual Backdoor Attack for Practical Applications
di: Song, Minkyoo, et al.
Pubblicazione: (2024)
di: Song, Minkyoo, et al.
Pubblicazione: (2024)
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
di: Kim, Jaehan, et al.
Pubblicazione: (2025)
di: Kim, Jaehan, et al.
Pubblicazione: (2025)
Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizers
di: Jang, Eugene, et al.
Pubblicazione: (2024)
di: Jang, Eugene, et al.
Pubblicazione: (2024)
PassREfinder-FL: Privacy-Preserving Credential Stuffing Risk Prediction via Graph-Based Federated Learning for Representing Password Reuse between Websites
di: Kim, Jaehan, et al.
Pubblicazione: (2025)
di: Kim, Jaehan, et al.
Pubblicazione: (2025)
Subgraph Reconstruction Attacks on Graph RAG Deployments with Practical Defenses
di: Song, Minkyoo, et al.
Pubblicazione: (2026)
di: Song, Minkyoo, et al.
Pubblicazione: (2026)
Detection of Illicit Content on Online Marketplaces using Large Language Models
di: Tran, Quoc Khoa, et al.
Pubblicazione: (2026)
di: Tran, Quoc Khoa, et al.
Pubblicazione: (2026)
Personalized Real-time Jargon Support for Online Meetings
di: Song, Yifan, et al.
Pubblicazione: (2025)
di: Song, Yifan, et al.
Pubblicazione: (2025)
IPS: In-Prompt Process Supervision for Short Video Content Moderation
di: Liu, Mingchao, et al.
Pubblicazione: (2024)
di: Liu, Mingchao, et al.
Pubblicazione: (2024)
Moderating Illicit Online Image Promotion for Unsafe User-Generated Content Games Using Large Vision-Language Models
di: Guo, Keyan, et al.
Pubblicazione: (2024)
di: Guo, Keyan, et al.
Pubblicazione: (2024)
Distantly Supervised Morpho-Syntactic Model for Relation Extraction
di: Gutehrlé, Nicolas, et al.
Pubblicazione: (2024)
di: Gutehrlé, Nicolas, et al.
Pubblicazione: (2024)
Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation
di: Jang, Won Seok, et al.
Pubblicazione: (2025)
di: Jang, Won Seok, et al.
Pubblicazione: (2025)
ExpGuard: LLM Content Moderation in Specialized Domains
di: Choi, Minseok, et al.
Pubblicazione: (2026)
di: Choi, Minseok, et al.
Pubblicazione: (2026)
Large Language Model-based Role-Playing for Personalized Medical Jargon Extraction
di: Lim, Jung Hoon, et al.
Pubblicazione: (2024)
di: Lim, Jung Hoon, et al.
Pubblicazione: (2024)
Cross-Lingual Transfer for Distantly Supervised and Low-resources Indonesian NER
di: Ikhwantri, Fariz
Pubblicazione: (2019)
di: Ikhwantri, Fariz
Pubblicazione: (2019)
Ignore Me But Don't Replace Me: Utilizing Non-Linguistic Elements for Pretraining on the Cybersecurity Domain
di: Jang, Eugene, et al.
Pubblicazione: (2024)
di: Jang, Eugene, et al.
Pubblicazione: (2024)
From Stars to Insights: Exploration and Implementation of Unified Sentiment Analysis with Distant Supervision
di: Li, Wenchang, et al.
Pubblicazione: (2023)
di: Li, Wenchang, et al.
Pubblicazione: (2023)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
di: Dinuta, Eduard Stefan, et al.
Pubblicazione: (2025)
di: Dinuta, Eduard Stefan, et al.
Pubblicazione: (2025)
Crossing Domains without Labels: Distant Supervision for Term Extraction
di: Senger, Elena, et al.
Pubblicazione: (2025)
di: Senger, Elena, et al.
Pubblicazione: (2025)
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning
di: Wu, Bohao, et al.
Pubblicazione: (2025)
di: Wu, Bohao, et al.
Pubblicazione: (2025)
Distantly-Supervised Joint Extraction with Noise-Robust Learning
di: Li, Yufei, et al.
Pubblicazione: (2023)
di: Li, Yufei, et al.
Pubblicazione: (2023)
Ideology-Based LLMs for Content Moderation
di: Civelli, Stefano, et al.
Pubblicazione: (2025)
di: Civelli, Stefano, et al.
Pubblicazione: (2025)
Constraint Multi-class Positive and Unlabeled Learning for Distantly Supervised Named Entity Recognition
di: Zhang, Yuzhe, et al.
Pubblicazione: (2025)
di: Zhang, Yuzhe, et al.
Pubblicazione: (2025)
Combining Distantly Supervised Models with In Context Learning for Monolingual and Cross-Lingual Relation Extraction
di: Rathore, Vipul, et al.
Pubblicazione: (2025)
di: Rathore, Vipul, et al.
Pubblicazione: (2025)
Can't say cant? Measuring and Reasoning of Dark Jargons in Large Language Models
di: Ji, Xu, et al.
Pubblicazione: (2024)
di: Ji, Xu, et al.
Pubblicazione: (2024)
Sentence Bag Graph Formulation for Biomedical Distant Supervision Relation Extraction
di: Zhang, Hao, et al.
Pubblicazione: (2023)
di: Zhang, Hao, et al.
Pubblicazione: (2023)
$PC^2$: Politically Controversial Content Generation via Jailbreaking Attacks on GPT-based Text-to-Image Models
di: Choi, Wonwoo, et al.
Pubblicazione: (2026)
di: Choi, Wonwoo, et al.
Pubblicazione: (2026)
AI Content Moderation in Therapy Conversations
di: Kim, Jiwon, et al.
Pubblicazione: (2026)
di: Kim, Jiwon, et al.
Pubblicazione: (2026)
DynClean: Training Dynamics-based Label Cleaning for Distantly-Supervised Named Entity Recognition
di: Zhang, Qi, et al.
Pubblicazione: (2025)
di: Zhang, Qi, et al.
Pubblicazione: (2025)
Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation
di: Li, Siyuan, et al.
Pubblicazione: (2026)
di: Li, Siyuan, et al.
Pubblicazione: (2026)
README: Bridging Medical Jargon and Lay Understanding for Patient Education through Data-Centric NLP
di: Yao, Zonghai, et al.
Pubblicazione: (2023)
di: Yao, Zonghai, et al.
Pubblicazione: (2023)
Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New Language
di: Plum, Alistair, et al.
Pubblicazione: (2024)
di: Plum, Alistair, et al.
Pubblicazione: (2024)
Re-Examine Distantly Supervised NER: A New Benchmark and a Simple Approach
di: Li, Yuepei, et al.
Pubblicazione: (2024)
di: Li, Yuepei, et al.
Pubblicazione: (2024)
When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs
di: Kim, Hanna, et al.
Pubblicazione: (2024)
di: Kim, Hanna, et al.
Pubblicazione: (2024)
DistALANER: Distantly Supervised Active Learning Augmented Named Entity Recognition in the Open Source Software Ecosystem
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
di: Banerjee, Somnath, et al.
Pubblicazione: (2024)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
di: Zheng, Jiangrui, et al.
Pubblicazione: (2023)
di: Zheng, Jiangrui, et al.
Pubblicazione: (2023)
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
di: Cao, Yang Trista, et al.
Pubblicazione: (2023)
di: Cao, Yang Trista, et al.
Pubblicazione: (2023)
LangBridge: Multilingual Reasoning Without Multilingual Supervision
di: Yoon, Dongkeun, et al.
Pubblicazione: (2024)
di: Yoon, Dongkeun, et al.
Pubblicazione: (2024)
BingoGuard: LLM Content Moderation Tools with Risk Levels
di: Yin, Fan, et al.
Pubblicazione: (2025)
di: Yin, Fan, et al.
Pubblicazione: (2025)
The Unappreciated Role of Intent in Algorithmic Moderation of Social Media Content
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm
di: Kim, Jaehan, et al.
Pubblicazione: (2024) -
Claim-Guided Textual Backdoor Attack for Practical Applications
di: Song, Minkyoo, et al.
Pubblicazione: (2024) -
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
di: Kim, Jaehan, et al.
Pubblicazione: (2025) -
Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizers
di: Jang, Eugene, et al.
Pubblicazione: (2024) -
PassREfinder-FL: Privacy-Preserving Credential Stuffing Risk Prediction via Graph-Based Federated Learning for Representing Password Reuse between Websites
di: Kim, Jaehan, et al.
Pubblicazione: (2025)