Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Yejin, Kim, Su-Hyeon, Jin, Hyundong, Kim, Dayoung, Kim, Yeonsoo, Han, Yo-Sub |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models
by: Lee, Yejin, et al.
Published: (2026)
by: Lee, Yejin, et al.
Published: (2026)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
Repairing Regex Vulnerabilities via Localization-Guided Instructions
by: Sung, Sicheol, et al.
Published: (2025)
by: Sung, Sicheol, et al.
Published: (2025)
Align-to-Distill: Trainable Attention Alignment for Knowledge Distillation in Neural Machine Translation
by: Jin, Heegon, et al.
Published: (2024)
by: Jin, Heegon, et al.
Published: (2024)
TCProF: Time-Complexity Prediction SSL Framework
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
Improving Commonsense Bias Classification by Mitigating the Influence of Demographic Terms
by: Lee, JinKyu, et al.
Published: (2024)
by: Lee, JinKyu, et al.
Published: (2024)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
by: Kim, Jaemin, et al.
Published: (2024)
by: Kim, Jaemin, et al.
Published: (2024)
Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For Perplexity
by: Seo, Yeongbin, et al.
Published: (2025)
by: Seo, Yeongbin, et al.
Published: (2025)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
by: Park, Seungcheol, et al.
Published: (2025)
by: Park, Seungcheol, et al.
Published: (2025)
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
by: Park, Seungcheol, et al.
Published: (2025)
by: Park, Seungcheol, et al.
Published: (2025)
"As Eastern Powers, I will veto." : An Investigation of Nation-level Bias of Large Language Models in International Relations
by: Choi, Jonghyeon, et al.
Published: (2025)
by: Choi, Jonghyeon, et al.
Published: (2025)
ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection
by: Delaval, Axel, et al.
Published: (2025)
by: Delaval, Axel, et al.
Published: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
by: Kim, Soyeon, et al.
Published: (2026)
by: Kim, Soyeon, et al.
Published: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
by: Kim, Heejun, et al.
Published: (2026)
by: Kim, Heejun, et al.
Published: (2026)
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
by: Seo, Yeongbin, et al.
Published: (2025)
by: Seo, Yeongbin, et al.
Published: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
by: Tu, Songjun, et al.
Published: (2026)
by: Tu, Songjun, et al.
Published: (2026)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
A Study on Bias Detection and Classification in Natural Language Processing
by: Evans, Ana Sofia, et al.
Published: (2024)
by: Evans, Ana Sofia, et al.
Published: (2024)
Topeax -- An Improved Clustering Topic Model with Density Peak Detection and Lexical-Semantic Term Importance
by: Kardos, Márton
Published: (2026)
by: Kardos, Márton
Published: (2026)
A Comprehensive Survey of Compression Algorithms for Language Models
by: Park, Seungcheol, et al.
Published: (2024)
by: Park, Seungcheol, et al.
Published: (2024)
Branching Narratives: Character Decision Points Detection
by: Tikhonov, Alexey
Published: (2024)
by: Tikhonov, Alexey
Published: (2024)
Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
by: Görge, Rebekka, et al.
Published: (2025)
by: Görge, Rebekka, et al.
Published: (2025)
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
by: Johnson, Warren, et al.
Published: (2026)
by: Johnson, Warren, et al.
Published: (2026)
WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents
by: Pan, Leyi, et al.
Published: (2024)
by: Pan, Leyi, et al.
Published: (2024)
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
by: Seo, Yeongbin, et al.
Published: (2024)
by: Seo, Yeongbin, et al.
Published: (2024)
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
by: Seo, Yeongbin, et al.
Published: (2025)
by: Seo, Yeongbin, et al.
Published: (2025)
Examining Linguistic Shifts in Academic Writing Before and After the Launch of ChatGPT: A Study on Preprint Papers
by: Bao, Tong, et al.
Published: (2025)
by: Bao, Tong, et al.
Published: (2025)
Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
by: Kim, Seongho, et al.
Published: (2024)
by: Kim, Seongho, et al.
Published: (2024)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
by: Nguyen, Minh Hoang, et al.
Published: (2025)
by: Nguyen, Minh Hoang, et al.
Published: (2025)
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
by: Sela, Omer
Published: (2026)
by: Sela, Omer
Published: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
by: Yang, Yibo
Published: (2025)
by: Yang, Yibo
Published: (2025)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
by: Marjanović, Sara Vera, et al.
Published: (2024)
by: Marjanović, Sara Vera, et al.
Published: (2024)
Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
by: An, Tao
Published: (2025)
by: An, Tao
Published: (2025)
One Agent to Serve All: a Lite-Adaptive Stylized AI Assistant for Millions of Multi-Style Official Accounts
by: Fan, Xingyu, et al.
Published: (2025)
by: Fan, Xingyu, et al.
Published: (2025)
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
by: Fang, Xi, et al.
Published: (2025)
by: Fang, Xi, et al.
Published: (2025)
TREX: Tokenizer Regression for Optimal Data Mixture
by: Won, Inho, et al.
Published: (2026)
by: Won, Inho, et al.
Published: (2026)
Language corpora for the Dutch medical domain
by: van Es, B.
Published: (2026)
by: van Es, B.
Published: (2026)
ADE: Adaptive Dictionary Embeddings -- Scaling Multi-Anchor Representations to Large Language Models
by: Demirci, Orhan, et al.
Published: (2026)
by: Demirci, Orhan, et al.
Published: (2026)
Similar Items
-
Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models
by: Lee, Yejin, et al.
Published: (2026) -
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
by: Lee, Yejin, et al.
Published: (2025) -
AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
by: Lee, Yejin, et al.
Published: (2025) -
Repairing Regex Vulnerabilities via Localization-Guided Instructions
by: Sung, Sicheol, et al.
Published: (2025) -
Align-to-Distill: Trainable Attention Alignment for Knowledge Distillation in Neural Machine Translation
by: Jin, Heegon, et al.
Published: (2024)