A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Machlovi, Naseem, Saleki, Maryam, Amin, Ruhul, Rahouti, Mohamed, Al-Maliki, Shawqi, Qadir, Junaid, Abdallah, Mohamed M., Al-Fuqaha, Ala |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Budget-Constrained Online Retrieval-Augmented Generation: The Chunk-as-a-Service Model
par: Al-Maliki, Shawqi, et autres
Publié: (2026)
par: Al-Maliki, Shawqi, et autres
Publié: (2026)
Addressing Data Distribution Shifts in Online Machine Learning Powered Smart City Applications Using Augmented Test-Time Adaptation
par: Al-Maliki, Shawqi, et autres
Publié: (2022)
par: Al-Maliki, Shawqi, et autres
Publié: (2022)
Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
par: Machlovi, Naseem, et autres
Publié: (2025)
par: Machlovi, Naseem, et autres
Publié: (2025)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
par: Mushtaq, Abdullah, et autres
Publié: (2025)
par: Mushtaq, Abdullah, et autres
Publié: (2025)
Large Language Model Enhanced Particle Swarm Optimization for Hyperparameter Tuning for Deep Learning Models
par: Hameed, Saad, et autres
Publié: (2025)
par: Hameed, Saad, et autres
Publié: (2025)
Efficient Data Labeling and Optimal Device Scheduling in HWNs Using Clustered Federated Semi-Supervised Learning
par: Hamood, Moqbel, et autres
Publié: (2024)
par: Hamood, Moqbel, et autres
Publié: (2024)
Safeguarding connected autonomous vehicle communication: Protocols, intra- and inter-vehicular attacks and defenses
par: Aledhari, Mohammed, et autres
Publié: (2025)
par: Aledhari, Mohammed, et autres
Publié: (2025)
TOSHFA: A Mobile VR-Based System for Pose-Guided Exercise Rehabilitation for Low Back Pain
par: Mohamed, Amin, et autres
Publié: (2026)
par: Mohamed, Amin, et autres
Publié: (2026)
BioEnvSense: A Human-Centred Security Framework for Preventing Behaviour-Driven Cyber Incidents
par: Ta, Duy Anh, et autres
Publié: (2026)
par: Ta, Duy Anh, et autres
Publié: (2026)
Charging Ahead: A Hierarchical Adversarial Framework for Counteracting Advanced Cyber Threats in EV Charging Stations
par: Al-Mehdhar, Mohammed, et autres
Publié: (2024)
par: Al-Mehdhar, Mohammed, et autres
Publié: (2024)
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
par: Mehta, Manisha, et autres
Publié: (2025)
par: Mehta, Manisha, et autres
Publié: (2025)
Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant
par: Matveyenko, Joseph, et autres
Publié: (2025)
par: Matveyenko, Joseph, et autres
Publié: (2025)
PersoPilot: An Adaptive AI-Copilot for Transparent Contextualized Persona Classification and Personalized Response Generation
par: Afzoon, Saleh, et autres
Publié: (2026)
par: Afzoon, Saleh, et autres
Publié: (2026)
MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
par: Jing, Haifeng, et autres
Publié: (2025)
par: Jing, Haifeng, et autres
Publié: (2025)
Composable Prompting Workspaces for Creative Writing: Exploration and Iteration Using Dynamic Widgets
par: Amin, Rifat Mehreen, et autres
Publié: (2025)
par: Amin, Rifat Mehreen, et autres
Publié: (2025)
PromptCanvas: Composable Prompting Workspaces Using Dynamic Widgets for Exploration and Iteration in Creative Writing
par: Amin, Rifat Mehreen, et autres
Publié: (2025)
par: Amin, Rifat Mehreen, et autres
Publié: (2025)
Consistent Valid Physically-Realizable Adversarial Attack against Crowd-flow Prediction Models
par: Ali, Hassan, et autres
Publié: (2023)
par: Ali, Hassan, et autres
Publié: (2023)
Trusting the Search: Unraveling Human Trust in Health Information from Google and ChatGPT
par: Sun, Xin, et autres
Publié: (2024)
par: Sun, Xin, et autres
Publié: (2024)
Measuring Large Language Models Dependency: Validating the Arabic Version of the LLM-D12 Scale
par: AlShakhsi, Sameha, et autres
Publié: (2025)
par: AlShakhsi, Sameha, et autres
Publié: (2025)
Protege Effect for Behaviour Change: Does Teaching Digital Stress Solutions to Others Reduce One's Own?
par: Alshakhsi, Sameha, et autres
Publié: (2025)
par: Alshakhsi, Sameha, et autres
Publié: (2025)
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs
par: Tytarenko, Stepan, et autres
Publié: (2024)
par: Tytarenko, Stepan, et autres
Publié: (2024)
Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription
par: Alkaeed, Mahdi, et autres
Publié: (2025)
par: Alkaeed, Mahdi, et autres
Publié: (2025)
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation
par: Jin, Meiqing, et autres
Publié: (2025)
par: Jin, Meiqing, et autres
Publié: (2025)
LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?
par: Kebe, Gaoussou Youssouf, et autres
Publié: (2025)
par: Kebe, Gaoussou Youssouf, et autres
Publié: (2025)
(De)Noise: Moderating the Inconsistency Between Human Decision-Makers
par: Grgić-Hlača, Nina, et autres
Publié: (2024)
par: Grgić-Hlača, Nina, et autres
Publié: (2024)
Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being
par: Ameen, Mohd Ruhul, et autres
Publié: (2025)
par: Ameen, Mohd Ruhul, et autres
Publié: (2025)
Understanding Parents' Desires in Moderating Children's Interactions with GenAI Chatbots through LLM-Generated Probes
par: Driscoll, John, et autres
Publié: (2026)
par: Driscoll, John, et autres
Publié: (2026)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
par: Bouchekif, Abdessalam, et autres
Publié: (2026)
par: Bouchekif, Abdessalam, et autres
Publié: (2026)
GPT-5 vs Other LLMs in Long Short-Context Performance
par: Esmi, Nima, et autres
Publié: (2026)
par: Esmi, Nima, et autres
Publié: (2026)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
par: Afzoon, Saleh, et autres
Publié: (2026)
par: Afzoon, Saleh, et autres
Publié: (2026)
Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
par: Hartmann, David, et autres
Publié: (2025)
par: Hartmann, David, et autres
Publié: (2025)
SkySim: A ROS2-based Simulation Environment for Natural Language Control of Drone Swarms using Large Language Models
par: Shibu, Aditya, et autres
Publié: (2026)
par: Shibu, Aditya, et autres
Publié: (2026)
Empowering HWNs with Efficient Data Labeling: A Clustered Federated Semi-Supervised Learning Approach
par: Hamood, Moqbel, et autres
Publié: (2024)
par: Hamood, Moqbel, et autres
Publié: (2024)
GPT-4 Emulates Average-Human Emotional Cognition from a Third-Person Perspective
par: Tak, Ala N., et autres
Publié: (2024)
par: Tak, Ala N., et autres
Publié: (2024)
Therapy as an NLP Task: Psychologists' Comparison of LLMs and Human Peers in CBT
par: Iftikhar, Zainab, et autres
Publié: (2024)
par: Iftikhar, Zainab, et autres
Publié: (2024)
Multi-trait User Simulation with Adaptive Decoding for Conversational Task Assistants
par: Ferreira, Rafael, et autres
Publié: (2024)
par: Ferreira, Rafael, et autres
Publié: (2024)
Generative AI in Multimodal User Interfaces: Trends, Challenges, and Cross-Platform Adaptability
par: Bieniek, J., et autres
Publié: (2024)
par: Bieniek, J., et autres
Publié: (2024)
Personalizing Content Moderation on Social Media: User Perspectives on Moderation Choices, Interface Design, and Labor
par: Jhaver, Shagun, et autres
Publié: (2023)
par: Jhaver, Shagun, et autres
Publié: (2023)
Optimized Federated Multitask Learning in Mobile Edge Networks: A Hybrid Client Selection and Model Aggregation Approach
par: Hamood, Moqbel, et autres
Publié: (2024)
par: Hamood, Moqbel, et autres
Publié: (2024)
Same Voice, Different Lab: On the Homogenization of Frontier LLM Personalities
par: Krishna, Avinash, et autres
Publié: (2026)
par: Krishna, Avinash, et autres
Publié: (2026)
Documents similaires
-
Budget-Constrained Online Retrieval-Augmented Generation: The Chunk-as-a-Service Model
par: Al-Maliki, Shawqi, et autres
Publié: (2026) -
Addressing Data Distribution Shifts in Online Machine Learning Powered Smart City Applications Using Augmented Test-Time Adaptation
par: Al-Maliki, Shawqi, et autres
Publié: (2022) -
Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
par: Machlovi, Naseem, et autres
Publié: (2025) -
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
par: Mushtaq, Abdullah, et autres
Publié: (2025) -
Large Language Model Enhanced Particle Swarm Optimization for Hyperparameter Tuning for Deep Learning Models
par: Hameed, Saad, et autres
Publié: (2025)