STAND-Guard: A Small Task-Adaptive Content Moderation Model
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Minjia, Lin, Pingping, Cai, Siqi, An, Shengnan, Ma, Shengjie, Lin, Zeqi, Huang, Congrui, Xu, Bixiong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compositional API Recommendation for Library-Oriented Code Generation
by: Ma, Zexiong, et al.
Published: (2024)
by: Ma, Zexiong, et al.
Published: (2024)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
by: Liu, Hongfu, et al.
Published: (2024)
by: Liu, Hongfu, et al.
Published: (2024)
ExpGuard: LLM Content Moderation in Specialized Domains
by: Choi, Minseok, et al.
Published: (2026)
by: Choi, Minseok, et al.
Published: (2026)
Make Your LLM Fully Utilize the Context
by: An, Shengnan, et al.
Published: (2024)
by: An, Shengnan, et al.
Published: (2024)
BingoGuard: LLM Content Moderation Tools with Risk Levels
by: Yin, Fan, et al.
Published: (2025)
by: Yin, Fan, et al.
Published: (2025)
Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
by: Wang, Zheng, et al.
Published: (2024)
by: Wang, Zheng, et al.
Published: (2024)
Dehallucinating Parallel Context Extension for Retrieval-Augmented Generation
by: Ma, Zexiong, et al.
Published: (2024)
by: Ma, Zexiong, et al.
Published: (2024)
Learning From Mistakes Makes LLM Better Reasoner
by: An, Shengnan, et al.
Published: (2023)
by: An, Shengnan, et al.
Published: (2023)
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
by: Jha, Prince, et al.
Published: (2024)
by: Jha, Prince, et al.
Published: (2024)
SLM-Mod: Small Language Models Surpass LLMs at Content Moderation
by: Zhan, Xianyang, et al.
Published: (2024)
by: Zhan, Xianyang, et al.
Published: (2024)
LionGuard: Building a Contextualized Moderation Classifier to Tackle Localized Unsafe Content
by: Foo, Jessica, et al.
Published: (2024)
by: Foo, Jessica, et al.
Published: (2024)
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
by: Tan, Leanne, et al.
Published: (2025)
by: Tan, Leanne, et al.
Published: (2025)
MTikGuard System: A Transformer-Based Multimodal System for Child-Safe Content Moderation on TikTok
by: Nguyen, Dat Thanh, et al.
Published: (2025)
by: Nguyen, Dat Thanh, et al.
Published: (2025)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)
by: Elesedy, Hayder, et al.
Published: (2024)
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
by: Fatehkia, Masoomali, et al.
Published: (2025)
by: Fatehkia, Masoomali, et al.
Published: (2025)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026)
by: Wróbel, Krzysztof, et al.
Published: (2026)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024)
by: Han, Seungju, et al.
Published: (2024)
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
by: Kumar, Priyanshu, et al.
Published: (2025)
by: Kumar, Priyanshu, et al.
Published: (2025)
Legilimens: Practical and Unified Content Moderation for Large Language Model Services
by: Wu, Jialin, et al.
Published: (2024)
by: Wu, Jialin, et al.
Published: (2024)
Short-PHD: Detecting Short LLM-generated Text with Topological Data Analysis After Off-topic Content Insertion
by: Wei, Dongjun, et al.
Published: (2025)
by: Wei, Dongjun, et al.
Published: (2025)
AI Content Moderation in Therapy Conversations
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Self-Guard: Empower the LLM to Safeguard Itself
by: Wang, Zezhong, et al.
Published: (2023)
by: Wang, Zezhong, et al.
Published: (2023)
Ideology-Based LLMs for Content Moderation
by: Civelli, Stefano, et al.
Published: (2025)
by: Civelli, Stefano, et al.
Published: (2025)
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
by: Ghosh, Shaona, et al.
Published: (2024)
by: Ghosh, Shaona, et al.
Published: (2024)
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
by: He, Yinghui, et al.
Published: (2025)
by: He, Yinghui, et al.
Published: (2025)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
by: Ge, Suyu, et al.
Published: (2023)
by: Ge, Suyu, et al.
Published: (2023)
SAMoRA: Semantic-Aware Mixture of LoRA Experts for Task-Adaptive Learning
by: Shi, Boyan, et al.
Published: (2026)
by: Shi, Boyan, et al.
Published: (2026)
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
by: Yu, Jiaqi, et al.
Published: (2026)
by: Yu, Jiaqi, et al.
Published: (2026)
Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
Class-RAG: Real-Time Content Moderation with Retrieval Augmented Generation
by: Chen, Jianfa, et al.
Published: (2024)
by: Chen, Jianfa, et al.
Published: (2024)
HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
by: Bergeron, Loris, et al.
Published: (2025)
by: Bergeron, Loris, et al.
Published: (2025)
The Unappreciated Role of Intent in Algorithmic Moderation of Social Media Content
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Measuring Stereotype and Deviation Biases in Large Language Models
by: Wang, Daniel, et al.
Published: (2025)
by: Wang, Daniel, et al.
Published: (2025)
ClickGuard: A Trustworthy Adaptive Fusion Framework for Clickbait Detection
by: Dhiman, Chhavi, et al.
Published: (2026)
by: Dhiman, Chhavi, et al.
Published: (2026)
Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards
by: Lyu, Yuanjie, et al.
Published: (2026)
by: Lyu, Yuanjie, et al.
Published: (2026)
ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information
by: Hui, Zheng, et al.
Published: (2024)
by: Hui, Zheng, et al.
Published: (2024)
A General Method for Detecting Information Generated by Large Language Models
by: Mao, Minjia, et al.
Published: (2025)
by: Mao, Minjia, et al.
Published: (2025)
PaperHelper: Knowledge-Based LLM QA Paper Reading Assistant
by: Yin, Congrui, et al.
Published: (2025)
by: Yin, Congrui, et al.
Published: (2025)
IAPT: Instruction-Aware Prompt Tuning for Large Language Models
by: Zhu, Wei, et al.
Published: (2024)
by: Zhu, Wei, et al.
Published: (2024)
Parrot Mind: Towards Explaining the Complex Task Reasoning of Pretrained Large Language Models with Template-Content Structure
by: Yang, Haotong, et al.
Published: (2023)
by: Yang, Haotong, et al.
Published: (2023)
Similar Items
-
Compositional API Recommendation for Library-Oriented Code Generation
by: Ma, Zexiong, et al.
Published: (2024) -
On Calibration of LLM-based Guard Models for Reliable Content Moderation
by: Liu, Hongfu, et al.
Published: (2024) -
ExpGuard: LLM Content Moderation in Specialized Domains
by: Choi, Minseok, et al.
Published: (2026) -
Make Your LLM Fully Utilize the Context
by: An, Shengnan, et al.
Published: (2024) -
BingoGuard: LLM Content Moderation Tools with Risk Levels
by: Yin, Fan, et al.
Published: (2025)