PAM: Training Policy-Aligned Moderation Filters at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Fatehkia, Masoomali, Altinisik, Enes, Osman, Mohamed, Sencar, Husrev Taha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
by: Fatehkia, Masoomali, et al.
Published: (2025)
by: Fatehkia, Masoomali, et al.
Published: (2025)
Tool Calling for Arabic LLMs: Data Strategies and Instruction Tuning
by: Ersoy, Asim, et al.
Published: (2025)
by: Ersoy, Asim, et al.
Published: (2025)
Explaining the role of Intrinsic Dimensionality in Adversarial Training
by: Altinisik, Enes, et al.
Published: (2024)
by: Altinisik, Enes, et al.
Published: (2024)
Do I Really Know? Learning Factual Self-Verification for Hallucination Reduction
by: Altinisik, Enes, et al.
Published: (2026)
by: Altinisik, Enes, et al.
Published: (2026)
T-RAG: Lessons from the LLM Trenches
by: Fatehkia, Masoomali, et al.
Published: (2024)
by: Fatehkia, Masoomali, et al.
Published: (2024)
Multimedia Forensics
by: Husrev Taha Sencar
by: Husrev Taha Sencar
There Is More to Refusal in Large Language Models than a Single Direction
by: Joad, Faaiz, et al.
Published: (2026)
by: Joad, Faaiz, et al.
Published: (2026)
From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction
by: Lekssays, Ahmed, et al.
Published: (2025)
by: Lekssays, Ahmed, et al.
Published: (2025)
Fanar 2.0: Arabic Generative AI Stack
by: FANAR TEAM, et al.
Published: (2026)
by: FANAR TEAM, et al.
Published: (2026)
TechniqueRAG: Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
by: Lekssays, Ahmed, et al.
Published: (2025)
by: Lekssays, Ahmed, et al.
Published: (2025)
Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling
by: Li, Anqi, et al.
Published: (2025)
by: Li, Anqi, et al.
Published: (2025)
Fanar: An Arabic-Centric Multimodal Generative AI Platform
by: Fanar Team, et al.
Published: (2025)
by: Fanar Team, et al.
Published: (2025)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
by: Zheng, Jiangrui, et al.
Published: (2023)
by: Zheng, Jiangrui, et al.
Published: (2023)
TSLFormer: A Lightweight Transformer Model for Turkish Sign Language Recognition Using Skeletal Landmarks
by: Ertürk, Kutay, et al.
Published: (2025)
by: Ertürk, Kutay, et al.
Published: (2025)
Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents
by: Wang, Minzheng, et al.
Published: (2026)
by: Wang, Minzheng, et al.
Published: (2026)
Semantic Ranking for Automated Adversarial Technique Annotation in Security Text
by: Kumarasinghe, Udesh, et al.
Published: (2024)
by: Kumarasinghe, Udesh, et al.
Published: (2024)
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
by: Singh, Ishneet Sukhvinder, et al.
Published: (2024)
by: Singh, Ishneet Sukhvinder, et al.
Published: (2024)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
by: Miyamoto, Sora, et al.
Published: (2026)
by: Miyamoto, Sora, et al.
Published: (2026)
Aligning Teacher with Student Preferences for Tailored Training Data Generation
by: Liu, Yantao, et al.
Published: (2024)
by: Liu, Yantao, et al.
Published: (2024)
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning
by: Wang, Daoyu, et al.
Published: (2026)
by: Wang, Daoyu, et al.
Published: (2026)
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
by: Wang, Yan, et al.
Published: (2026)
by: Wang, Yan, et al.
Published: (2026)
A Multi-Dimensional Audit of Politically Aligned Large Language Models
by: Korver, Lisa, et al.
Published: (2026)
by: Korver, Lisa, et al.
Published: (2026)
ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws
by: Li, Ruihang, et al.
Published: (2024)
by: Li, Ruihang, et al.
Published: (2024)
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
by: Alrashed, Sultan, et al.
Published: (2025)
by: Alrashed, Sultan, et al.
Published: (2025)
Align then Train: Efficient Retrieval Adapter Learning
by: Maekawa, Seiji, et al.
Published: (2026)
by: Maekawa, Seiji, et al.
Published: (2026)
SER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition
by: Osman, Mohamed, et al.
Published: (2024)
by: Osman, Mohamed, et al.
Published: (2024)
Aligning Neural Machine Translation Models: Human Feedback in Training and Inference
by: Ramos, Miguel Moura, et al.
Published: (2023)
by: Ramos, Miguel Moura, et al.
Published: (2023)
ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection
by: Li, Boyang, et al.
Published: (2026)
by: Li, Boyang, et al.
Published: (2026)
SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant Hostility
by: Su, Xuanyu, et al.
Published: (2026)
by: Su, Xuanyu, et al.
Published: (2026)
SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling
by: Sun, Quanen, et al.
Published: (2026)
by: Sun, Quanen, et al.
Published: (2026)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
Prior Constraints-based Reward Model Training for Aligning Large Language Models
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
A Comprehensive Survey of Text Classification Techniques and Their Research Applications: Observational and Experimental Insights
by: Taha, Kamal, et al.
Published: (2024)
by: Taha, Kamal, et al.
Published: (2024)
CLASS-IT: Conversational and Lecture-Aligned Small-Scale Instruction Tuning for BabyLMs
by: Capone, Luca, et al.
Published: (2025)
by: Capone, Luca, et al.
Published: (2025)
Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
by: Si, Shuzheng, et al.
Published: (2025)
by: Si, Shuzheng, et al.
Published: (2025)
SyncThink: A Training-Free Strategy to Align Inference Termination with Reasoning Saturation
by: Li, Gengyang, et al.
Published: (2026)
by: Li, Gengyang, et al.
Published: (2026)
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
by: Cheng, Jiale, et al.
Published: (2023)
by: Cheng, Jiale, et al.
Published: (2023)
Training-Free Group Relative Policy Optimization
by: Cai, Yuzheng, et al.
Published: (2025)
by: Cai, Yuzheng, et al.
Published: (2025)
A Rate-Distortion Framework for Summarization
by: Arda, Enes, et al.
Published: (2025)
by: Arda, Enes, et al.
Published: (2025)
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
by: Gajewska, Ewelina, et al.
Published: (2026)
by: Gajewska, Ewelina, et al.
Published: (2026)
Similar Items
-
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
by: Fatehkia, Masoomali, et al.
Published: (2025) -
Tool Calling for Arabic LLMs: Data Strategies and Instruction Tuning
by: Ersoy, Asim, et al.
Published: (2025) -
Explaining the role of Intrinsic Dimensionality in Adversarial Training
by: Altinisik, Enes, et al.
Published: (2024) -
Do I Really Know? Learning Factual Self-Verification for Hallucination Reduction
by: Altinisik, Enes, et al.
Published: (2026) -
T-RAG: Lessons from the LLM Trenches
by: Fatehkia, Masoomali, et al.
Published: (2024)