Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
Fuente:
arXiv
Saved in:
| Main Authors: | Mehta, Manisha, Giunchiglia, Fausto |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
"I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated Dialogue
by: Taheri, Atieh, et al.
Published: (2026)
by: Taheri, Atieh, et al.
Published: (2026)
Collective Constitutional AI: Aligning a Language Model with Public Input
by: Huang, Saffron, et al.
Published: (2024)
by: Huang, Saffron, et al.
Published: (2024)
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
by: Du, Yishan, et al.
Published: (2025)
by: Du, Yishan, et al.
Published: (2025)
Generative UI as an Accessibility Bridge: Lessons from C2C E-Commerce
by: Ryskeldiev, Bektur
Published: (2026)
by: Ryskeldiev, Bektur
Published: (2026)
Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News
by: Loth, Alexander, et al.
Published: (2026)
by: Loth, Alexander, et al.
Published: (2026)
Chatbot Deployment Considerations for Application-Agnostic Human-Machine Dialogues
by: Rivas, Pablo, et al.
Published: (2025)
by: Rivas, Pablo, et al.
Published: (2025)
AI-induced sexual harassment: Investigating Contextual Characteristics and User Reactions of Sexual Harassment by a Companion Chatbot
by: Mohammad, et al.
Published: (2025)
by: Mohammad, et al.
Published: (2025)
Textual Entailment is not a Better Bias Metric than Token Probability
by: Felkner, Virginia K., et al.
Published: (2025)
by: Felkner, Virginia K., et al.
Published: (2025)
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
by: Felkner, Virginia K., et al.
Published: (2024)
by: Felkner, Virginia K., et al.
Published: (2024)
How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses
by: Gupta, Ishan, et al.
Published: (2026)
by: Gupta, Ishan, et al.
Published: (2026)
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026)
by: Maheshwari, Aditya, et al.
Published: (2026)
A Contextual Help Browser Extension to Assist Digital Illiterate Internet Users
by: Koutsiaris, Christos
Published: (2026)
by: Koutsiaris, Christos
Published: (2026)
Generative Confidants: How do People Experience Trust in Emotional Support from Generative AI?
by: Volpato, Riccardo, et al.
Published: (2026)
by: Volpato, Riccardo, et al.
Published: (2026)
LLMs as Educational Analysts: Transforming Multimodal Data Traces into Actionable Reading Assessment Reports
by: Davalos, Eduardo, et al.
Published: (2025)
by: Davalos, Eduardo, et al.
Published: (2025)
Using a cognitive architecture to consider antiBlackness in design and development of AI systems
by: Dancy, Christopher L.
Published: (2022)
by: Dancy, Christopher L.
Published: (2022)
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
by: Kuric, Eduard, et al.
Published: (2026)
by: Kuric, Eduard, et al.
Published: (2026)
LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight
by: Wachowiak, Lennart, et al.
Published: (2026)
by: Wachowiak, Lennart, et al.
Published: (2026)
LLM-Driven Accessible Interface: A Model-Based Approach
by: Jerry, Blessing, et al.
Published: (2026)
by: Jerry, Blessing, et al.
Published: (2026)
Understanding Parents' Desires in Moderating Children's Interactions with GenAI Chatbots through LLM-Generated Probes
by: Driscoll, John, et al.
Published: (2026)
by: Driscoll, John, et al.
Published: (2026)
Rejected Dialects: Biases Against African American Language in Reward Models
by: Mire, Joel, et al.
Published: (2025)
by: Mire, Joel, et al.
Published: (2025)
Characterizing Resource Sharing Practices on Underground Internet Forum Synthetic Non-Consensual Intimate Image Content Creation Communities
by: Medeiros, Bernardo B. P., et al.
Published: (2026)
by: Medeiros, Bernardo B. P., et al.
Published: (2026)
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
by: Alberts, Lize, et al.
Published: (2024)
by: Alberts, Lize, et al.
Published: (2024)
Will AI shape the way we speak? The emerging sociolinguistic influence of synthetic voices
by: Székely, Éva, et al.
Published: (2025)
by: Székely, Éva, et al.
Published: (2025)
The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims
by: Kelly, Matthew
Published: (2025)
by: Kelly, Matthew
Published: (2025)
AI-driven formative assessment and adaptive learning in data-science education: Evaluating an LLM-powered virtual teaching assistant
by: Anaroua, Fadjimata I, et al.
Published: (2025)
by: Anaroua, Fadjimata I, et al.
Published: (2025)
Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
by: Sami, K. M. Jubair, et al.
Published: (2026)
by: Sami, K. M. Jubair, et al.
Published: (2026)
How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)
by: Himmelreich, Johannes
Published: (2026)
by: Himmelreich, Johannes
Published: (2026)
Identifying Features that Shape Perceived Consciousness in Large Language Model-based AI: A Quantitative Study of Human Responses
by: Kang, Bongsu, et al.
Published: (2025)
by: Kang, Bongsu, et al.
Published: (2025)
Evaluation of Hate Speech Detection Using Large Language Models and Geographical Contextualization
by: Zahid, Anwar Hossain, et al.
Published: (2025)
by: Zahid, Anwar Hossain, et al.
Published: (2025)
How AI Systems Think About Education: Analyzing Latent Preference Patterns in Large Language Models
by: Autenrieth, Daniel
Published: (2026)
by: Autenrieth, Daniel
Published: (2026)
Know Your Author: Does the AI Penalty Hold in Short Fiction?
by: Todasco, Michael, et al.
Published: (2026)
by: Todasco, Michael, et al.
Published: (2026)
Playing Games with My Heart: An Evaluation of AI Companion Apps
by: Rauh, Maribeth, et al.
Published: (2026)
by: Rauh, Maribeth, et al.
Published: (2026)
Evaluating Epistemic Guardrails in AI Reading Assistants: A Behavioral Audit of a Minimal Prototype
by: Agustin, Matthew Christian
Published: (2026)
by: Agustin, Matthew Christian
Published: (2026)
Towards Democratized Flood Risk Management: An Advanced AI Assistant Enabled by GPT-4 for Enhanced Interpretability and Public Engagement
by: Martelo, Rafaela, et al.
Published: (2024)
by: Martelo, Rafaela, et al.
Published: (2024)
Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
by: Navarro, Karla Felix, et al.
Published: (2026)
by: Navarro, Karla Felix, et al.
Published: (2026)
AdaptiveCoPilot: Design and Testing of a NeuroAdaptive LLM Cockpit Guidance System in both Novice and Expert Pilots
by: Wen, Shaoyue, et al.
Published: (2025)
by: Wen, Shaoyue, et al.
Published: (2025)
The Reliance Negotiation Framework: A Dynamic Process Model of Student LLM Engagement in Academic Writing
by: Hossain, Shahin
Published: (2026)
by: Hossain, Shahin
Published: (2026)
"The Data Says Otherwise"-Towards Automated Fact-checking and Communication of Data Claims
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
Aurora: Neuro-Symbolic AI Driven Advising Agent
by: Lugones, Lorena Amanda Quincoso, et al.
Published: (2026)
by: Lugones, Lorena Amanda Quincoso, et al.
Published: (2026)
Talk, Listen, Connect: How Humans and AI Evaluate Empathy in Responses to Emotionally Charged Narratives
by: Roshanaei, Mahnaz, et al.
Published: (2024)
by: Roshanaei, Mahnaz, et al.
Published: (2024)
Similar Items
-
"I followed what felt right, not what I was told": Autonomy, Coaching, and Recognizing Bias Through AI-Mediated Dialogue
by: Taheri, Atieh, et al.
Published: (2026) -
Collective Constitutional AI: Aligning a Language Model with Public Input
by: Huang, Saffron, et al.
Published: (2024) -
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
by: Du, Yishan, et al.
Published: (2025) -
Generative UI as an Accessibility Bridge: Lessons from C2C E-Commerce
by: Ryskeldiev, Bektur
Published: (2026) -
Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News
by: Loth, Alexander, et al.
Published: (2026)