ExpGuard: LLM Content Moderation in Specialized Domains
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Choi, Minseok, Kim, Dongjin, Yang, Seungbin, Kim, Subin, Kwak, Youngjun, Oh, Juyoung, Choo, Jaegul, Son, Jungmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
von: Lee, Yunseung, et al.
Veröffentlicht: (2026)
von: Lee, Yunseung, et al.
Veröffentlicht: (2026)
LiveWeb-IE: A Benchmark For Online Web Information Extraction
von: Yang, Seungbin, et al.
Veröffentlicht: (2026)
von: Yang, Seungbin, et al.
Veröffentlicht: (2026)
Retrieve Only Relevant Tables Whether Few or Many: Adaptive Table Retrieval Method
von: Kim, Taehee, et al.
Veröffentlicht: (2026)
von: Kim, Taehee, et al.
Veröffentlicht: (2026)
Can Tool-augmented Large Language Models be Aware of Incomplete Conditions?
von: Yang, Seungbin, et al.
Veröffentlicht: (2024)
von: Yang, Seungbin, et al.
Veröffentlicht: (2024)
Cross-Lingual Unlearning of Selective Knowledge in Multilingual Language Models
von: Choi, Minseok, et al.
Veröffentlicht: (2024)
von: Choi, Minseok, et al.
Veröffentlicht: (2024)
Opt-Out: Investigating Entity-Level Unlearning for Large Language Models via Optimal Transport
von: Choi, Minseok, et al.
Veröffentlicht: (2024)
von: Choi, Minseok, et al.
Veröffentlicht: (2024)
Protecting Privacy Through Approximating Optimal Parameters for Sequence Unlearning in Language Models
von: Lee, Dohyun, et al.
Veröffentlicht: (2024)
von: Lee, Dohyun, et al.
Veröffentlicht: (2024)
Breaking Chains: Unraveling the Links in Multi-Hop Knowledge Unlearning
von: Choi, Minseok, et al.
Veröffentlicht: (2024)
von: Choi, Minseok, et al.
Veröffentlicht: (2024)
PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
BingoGuard: LLM Content Moderation Tools with Risk Levels
von: Yin, Fan, et al.
Veröffentlicht: (2025)
von: Yin, Fan, et al.
Veröffentlicht: (2025)
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
von: Son, Youngjun, et al.
Veröffentlicht: (2025)
von: Son, Youngjun, et al.
Veröffentlicht: (2025)
FENCE: A Financial and Multimodal Jailbreak Detection Dataset
von: Kim, Mirae, et al.
Veröffentlicht: (2026)
von: Kim, Mirae, et al.
Veröffentlicht: (2026)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
von: Jha, Prince, et al.
Veröffentlicht: (2024)
von: Jha, Prince, et al.
Veröffentlicht: (2024)
When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models
von: Park, Jungwon, et al.
Veröffentlicht: (2026)
von: Park, Jungwon, et al.
Veröffentlicht: (2026)
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
STAND-Guard: A Small Task-Adaptive Content Moderation Model
von: Wang, Minjia, et al.
Veröffentlicht: (2024)
von: Wang, Minjia, et al.
Veröffentlicht: (2024)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
von: Wróbel, Krzysztof, et al.
Veröffentlicht: (2026)
von: Wróbel, Krzysztof, et al.
Veröffentlicht: (2026)
The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators
von: Jeong, Hawon, et al.
Veröffentlicht: (2024)
von: Jeong, Hawon, et al.
Veröffentlicht: (2024)
Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information
von: Cho, Hojun, et al.
Veröffentlicht: (2025)
von: Cho, Hojun, et al.
Veröffentlicht: (2025)
Aligning Extraction and Generation for Robust Retrieval-Augmented Generation
von: Song, Hwanjun, et al.
Veröffentlicht: (2025)
von: Song, Hwanjun, et al.
Veröffentlicht: (2025)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
Efficient Terminology Integration for LLM-based Translation in Specialized Domains
von: Kim, Sejoon, et al.
Veröffentlicht: (2024)
von: Kim, Sejoon, et al.
Veröffentlicht: (2024)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
Scaling Up LLM Reviews for Google Ads Content Moderation
von: Qiao, Wei, et al.
Veröffentlicht: (2024)
von: Qiao, Wei, et al.
Veröffentlicht: (2024)
Federated Learning for Face Recognition via Intra-subject Self-supervised Learning
von: Kim, Hansol, et al.
Veröffentlicht: (2024)
von: Kim, Hansol, et al.
Veröffentlicht: (2024)
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
von: Foo, Jessica, et al.
Veröffentlicht: (2024)
von: Foo, Jessica, et al.
Veröffentlicht: (2024)
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
von: Tan, Leanne, et al.
Veröffentlicht: (2025)
von: Tan, Leanne, et al.
Veröffentlicht: (2025)
VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model
von: Choi, Junhyuk, et al.
Veröffentlicht: (2025)
von: Choi, Junhyuk, et al.
Veröffentlicht: (2025)
LaDiMo: Layer-wise Distillation Inspired MoEfier
von: Kim, Sungyoon, et al.
Veröffentlicht: (2024)
von: Kim, Sungyoon, et al.
Veröffentlicht: (2024)
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
von: Kwak, Beong-woo, et al.
Veröffentlicht: (2025)
von: Kwak, Beong-woo, et al.
Veröffentlicht: (2025)
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
von: Yu, Seunguk, et al.
Veröffentlicht: (2025)
von: Yu, Seunguk, et al.
Veröffentlicht: (2025)
Enhancing Intrinsic Features for Debiasing via Investigating Class-Discerning Common Attributes in Bias-Contrastive Pair
von: Park, Jeonghoon, et al.
Veröffentlicht: (2024)
von: Park, Jeonghoon, et al.
Veröffentlicht: (2024)
Exploring In-context Example Generation for Machine Translation
von: Lee, Dohyun, et al.
Veröffentlicht: (2025)
von: Lee, Dohyun, et al.
Veröffentlicht: (2025)
Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
von: Park, Cheonbok, et al.
Veröffentlicht: (2025)
von: Park, Cheonbok, et al.
Veröffentlicht: (2025)
MTikGuard System: A Transformer-Based Multimodal System for Child-Safe Content Moderation on TikTok
von: Nguyen, Dat Thanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Dat Thanh, et al.
Veröffentlicht: (2025)
Bones Can't Be Triangles: Accurate and Efficient Vertebrae Keypoint Estimation through Collaborative Error Revision
von: Kim, Jinhee, et al.
Veröffentlicht: (2024)
von: Kim, Jinhee, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BankMathBench: A Benchmark for Numerical Reasoning in Banking Scenarios
von: Lee, Yunseung, et al.
Veröffentlicht: (2026) -
LiveWeb-IE: A Benchmark For Online Web Information Extraction
von: Yang, Seungbin, et al.
Veröffentlicht: (2026) -
Retrieve Only Relevant Tables Whether Few or Many: Adaptive Table Retrieval Method
von: Kim, Taehee, et al.
Veröffentlicht: (2026) -
Can Tool-augmented Large Language Models be Aware of Incomplete Conditions?
von: Yang, Seungbin, et al.
Veröffentlicht: (2024) -
Cross-Lingual Unlearning of Selective Knowledge in Multilingual Language Models
von: Choi, Minseok, et al.
Veröffentlicht: (2024)