A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
Fuente:
arXiv
Guardado en:
| Autores principales: | Machlovi, Naseem, Saleki, Maryam, Amin, Ruhul, Rahouti, Mohamed, Al-Maliki, Shawqi, Qadir, Junaid, Abdallah, Mohamed M., Al-Fuqaha, Ala |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Budget-Constrained Online Retrieval-Augmented Generation: The Chunk-as-a-Service Model
por: Al-Maliki, Shawqi, et al.
Publicado: (2026)
por: Al-Maliki, Shawqi, et al.
Publicado: (2026)
Addressing Data Distribution Shifts in Online Machine Learning Powered Smart City Applications Using Augmented Test-Time Adaptation
por: Al-Maliki, Shawqi, et al.
Publicado: (2022)
por: Al-Maliki, Shawqi, et al.
Publicado: (2022)
Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
por: Machlovi, Naseem, et al.
Publicado: (2025)
por: Machlovi, Naseem, et al.
Publicado: (2025)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
por: Mushtaq, Abdullah, et al.
Publicado: (2025)
por: Mushtaq, Abdullah, et al.
Publicado: (2025)
Large Language Model Enhanced Particle Swarm Optimization for Hyperparameter Tuning for Deep Learning Models
por: Hameed, Saad, et al.
Publicado: (2025)
por: Hameed, Saad, et al.
Publicado: (2025)
Efficient Data Labeling and Optimal Device Scheduling in HWNs Using Clustered Federated Semi-Supervised Learning
por: Hamood, Moqbel, et al.
Publicado: (2024)
por: Hamood, Moqbel, et al.
Publicado: (2024)
Safeguarding connected autonomous vehicle communication: Protocols, intra- and inter-vehicular attacks and defenses
por: Aledhari, Mohammed, et al.
Publicado: (2025)
por: Aledhari, Mohammed, et al.
Publicado: (2025)
TOSHFA: A Mobile VR-Based System for Pose-Guided Exercise Rehabilitation for Low Back Pain
por: Mohamed, Amin, et al.
Publicado: (2026)
por: Mohamed, Amin, et al.
Publicado: (2026)
BioEnvSense: A Human-Centred Security Framework for Preventing Behaviour-Driven Cyber Incidents
por: Ta, Duy Anh, et al.
Publicado: (2026)
por: Ta, Duy Anh, et al.
Publicado: (2026)
Charging Ahead: A Hierarchical Adversarial Framework for Counteracting Advanced Cyber Threats in EV Charging Stations
por: Al-Mehdhar, Mohammed, et al.
Publicado: (2024)
por: Al-Mehdhar, Mohammed, et al.
Publicado: (2024)
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
por: Mehta, Manisha, et al.
Publicado: (2025)
por: Mehta, Manisha, et al.
Publicado: (2025)
Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant
por: Matveyenko, Joseph, et al.
Publicado: (2025)
por: Matveyenko, Joseph, et al.
Publicado: (2025)
PersoPilot: An Adaptive AI-Copilot for Transparent Contextualized Persona Classification and Personalized Response Generation
por: Afzoon, Saleh, et al.
Publicado: (2026)
por: Afzoon, Saleh, et al.
Publicado: (2026)
MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
por: Jing, Haifeng, et al.
Publicado: (2025)
por: Jing, Haifeng, et al.
Publicado: (2025)
Composable Prompting Workspaces for Creative Writing: Exploration and Iteration Using Dynamic Widgets
por: Amin, Rifat Mehreen, et al.
Publicado: (2025)
por: Amin, Rifat Mehreen, et al.
Publicado: (2025)
PromptCanvas: Composable Prompting Workspaces Using Dynamic Widgets for Exploration and Iteration in Creative Writing
por: Amin, Rifat Mehreen, et al.
Publicado: (2025)
por: Amin, Rifat Mehreen, et al.
Publicado: (2025)
Consistent Valid Physically-Realizable Adversarial Attack against Crowd-flow Prediction Models
por: Ali, Hassan, et al.
Publicado: (2023)
por: Ali, Hassan, et al.
Publicado: (2023)
Trusting the Search: Unraveling Human Trust in Health Information from Google and ChatGPT
por: Sun, Xin, et al.
Publicado: (2024)
por: Sun, Xin, et al.
Publicado: (2024)
Measuring Large Language Models Dependency: Validating the Arabic Version of the LLM-D12 Scale
por: AlShakhsi, Sameha, et al.
Publicado: (2025)
por: AlShakhsi, Sameha, et al.
Publicado: (2025)
Protege Effect for Behaviour Change: Does Teaching Digital Stress Solutions to Others Reduce One's Own?
por: Alshakhsi, Sameha, et al.
Publicado: (2025)
por: Alshakhsi, Sameha, et al.
Publicado: (2025)
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs
por: Tytarenko, Stepan, et al.
Publicado: (2024)
por: Tytarenko, Stepan, et al.
Publicado: (2024)
Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription
por: Alkaeed, Mahdi, et al.
Publicado: (2025)
por: Alkaeed, Mahdi, et al.
Publicado: (2025)
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation
por: Jin, Meiqing, et al.
Publicado: (2025)
por: Jin, Meiqing, et al.
Publicado: (2025)
LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?
por: Kebe, Gaoussou Youssouf, et al.
Publicado: (2025)
por: Kebe, Gaoussou Youssouf, et al.
Publicado: (2025)
(De)Noise: Moderating the Inconsistency Between Human Decision-Makers
por: Grgić-Hlača, Nina, et al.
Publicado: (2024)
por: Grgić-Hlača, Nina, et al.
Publicado: (2024)
Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being
por: Ameen, Mohd Ruhul, et al.
Publicado: (2025)
por: Ameen, Mohd Ruhul, et al.
Publicado: (2025)
Understanding Parents' Desires in Moderating Children's Interactions with GenAI Chatbots through LLM-Generated Probes
por: Driscoll, John, et al.
Publicado: (2026)
por: Driscoll, John, et al.
Publicado: (2026)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
por: Bouchekif, Abdessalam, et al.
Publicado: (2026)
por: Bouchekif, Abdessalam, et al.
Publicado: (2026)
GPT-5 vs Other LLMs in Long Short-Context Performance
por: Esmi, Nima, et al.
Publicado: (2026)
por: Esmi, Nima, et al.
Publicado: (2026)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
por: Afzoon, Saleh, et al.
Publicado: (2026)
por: Afzoon, Saleh, et al.
Publicado: (2026)
Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
por: Hartmann, David, et al.
Publicado: (2025)
por: Hartmann, David, et al.
Publicado: (2025)
SkySim: A ROS2-based Simulation Environment for Natural Language Control of Drone Swarms using Large Language Models
por: Shibu, Aditya, et al.
Publicado: (2026)
por: Shibu, Aditya, et al.
Publicado: (2026)
Empowering HWNs with Efficient Data Labeling: A Clustered Federated Semi-Supervised Learning Approach
por: Hamood, Moqbel, et al.
Publicado: (2024)
por: Hamood, Moqbel, et al.
Publicado: (2024)
GPT-4 Emulates Average-Human Emotional Cognition from a Third-Person Perspective
por: Tak, Ala N., et al.
Publicado: (2024)
por: Tak, Ala N., et al.
Publicado: (2024)
Therapy as an NLP Task: Psychologists' Comparison of LLMs and Human Peers in CBT
por: Iftikhar, Zainab, et al.
Publicado: (2024)
por: Iftikhar, Zainab, et al.
Publicado: (2024)
Multi-trait User Simulation with Adaptive Decoding for Conversational Task Assistants
por: Ferreira, Rafael, et al.
Publicado: (2024)
por: Ferreira, Rafael, et al.
Publicado: (2024)
Generative AI in Multimodal User Interfaces: Trends, Challenges, and Cross-Platform Adaptability
por: Bieniek, J., et al.
Publicado: (2024)
por: Bieniek, J., et al.
Publicado: (2024)
Personalizing Content Moderation on Social Media: User Perspectives on Moderation Choices, Interface Design, and Labor
por: Jhaver, Shagun, et al.
Publicado: (2023)
por: Jhaver, Shagun, et al.
Publicado: (2023)
Optimized Federated Multitask Learning in Mobile Edge Networks: A Hybrid Client Selection and Model Aggregation Approach
por: Hamood, Moqbel, et al.
Publicado: (2024)
por: Hamood, Moqbel, et al.
Publicado: (2024)
Same Voice, Different Lab: On the Homogenization of Frontier LLM Personalities
por: Krishna, Avinash, et al.
Publicado: (2026)
por: Krishna, Avinash, et al.
Publicado: (2026)
Ejemplares similares
-
Budget-Constrained Online Retrieval-Augmented Generation: The Chunk-as-a-Service Model
por: Al-Maliki, Shawqi, et al.
Publicado: (2026) -
Addressing Data Distribution Shifts in Online Machine Learning Powered Smart City Applications Using Augmented Test-Time Adaptation
por: Al-Maliki, Shawqi, et al.
Publicado: (2022) -
Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach
por: Machlovi, Naseem, et al.
Publicado: (2025) -
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
por: Mushtaq, Abdullah, et al.
Publicado: (2025) -
Large Language Model Enhanced Particle Swarm Optimization for Hyperparameter Tuning for Deep Learning Models
por: Hameed, Saad, et al.
Publicado: (2025)