Safe-Child-LLM: A Developmental Benchmark for Evaluating LLM Safety in Child-LLM Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Jiao, Junfeng, Afroogh, Saleh, Chen, Kevin, Murali, Abhejay, Atkinson, David, Dhurandhar, Amit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs and Childhood Safety: Identifying Risks and Proposing a Protection Framework for Safe Child-LLM Interaction
by: Jiao, Junfeng, et al.
Published: (2025)
by: Jiao, Junfeng, et al.
Published: (2025)
Evaluating LLM Safety Across Child Development Stages: A Simulated Agent Approach
by: Murali, Abhejay, et al.
Published: (2025)
by: Murali, Abhejay, et al.
Published: (2025)
LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models
by: Jiao, Junfeng, et al.
Published: (2025)
by: Jiao, Junfeng, et al.
Published: (2025)
LLM Harms: A Taxonomy and Discussion
by: Chen, Kevin, et al.
Published: (2025)
by: Chen, Kevin, et al.
Published: (2025)
AI Empathy Erodes Cognitive Autonomy in Younger Users
by: Jiao, Junfeng, et al.
Published: (2026)
by: Jiao, Junfeng, et al.
Published: (2026)
Generative AI and LLMs in Industry: A text-mining Analysis and Critical Evaluation of Guidelines and Policy Statements Across Fourteen Industrial Sectors
by: Jiao, Junfeng, et al.
Published: (2025)
by: Jiao, Junfeng, et al.
Published: (2025)
IGGA: A Dataset of Industrial Guidelines and Policy Statements for Generative AIs
by: Jiao, Junfeng, et al.
Published: (2025)
by: Jiao, Junfeng, et al.
Published: (2025)
AGGA: A Dataset of Academic Guidelines for Generative AI and Large Language Models
by: Jiao, Junfeng, et al.
Published: (2025)
by: Jiao, Junfeng, et al.
Published: (2025)
The global landscape of academic guidelines for generative AI and Large Language Models
by: Jiao, Junfeng, et al.
Published: (2024)
by: Jiao, Junfeng, et al.
Published: (2024)
Navigating LLM Ethics: Advancements, Challenges, and Future Directions
by: Jiao, Junfeng, et al.
Published: (2024)
by: Jiao, Junfeng, et al.
Published: (2024)
Evaluating the Effectiveness of OpenAI's Parental Control System
by: Ersoz, Kerem, et al.
Published: (2026)
by: Ersoz, Kerem, et al.
Published: (2026)
Intelligent Environmental Empathy (IEE): A new power and platform to fostering green obligation for climate peace and justice
by: Afroogh, Saleh, et al.
Published: (2024)
by: Afroogh, Saleh, et al.
Published: (2024)
Mapping out AI Functions in Intelligent Disaster (Mis)Management and AI-Caused Disasters
by: Pouresmaeil, Yasser, et al.
Published: (2025)
by: Pouresmaeil, Yasser, et al.
Published: (2025)
Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots
by: Park, Jihyung, et al.
Published: (2025)
by: Park, Jihyung, et al.
Published: (2025)
Applying LLM-Powered Virtual Humans to Child Interviews in Child-Centered Design
by: Li, Linshi, et al.
Published: (2025)
by: Li, Linshi, et al.
Published: (2025)
AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue
by: Park, Jihyung, et al.
Published: (2026)
by: Park, Jihyung, et al.
Published: (2026)
A Task-Driven Human-AI Collaboration: When to Automate, When to Collaborate, When to Challenge
by: Afroogh, Saleh, et al.
Published: (2025)
by: Afroogh, Saleh, et al.
Published: (2025)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
When Trust is Zero Sum: Automation Threat to Epistemic Agency
by: Malone, Emmie, et al.
Published: (2024)
by: Malone, Emmie, et al.
Published: (2024)
LLM Safety for Children
by: Rath, Prasanjit, et al.
Published: (2025)
by: Rath, Prasanjit, et al.
Published: (2025)
Dean of LLM Tutors: Exploring Comprehensive and Automated Evaluation of LLM-generated Educational Feedback via LLM Feedback Evaluators
by: Qian, Keyang, et al.
Published: (2025)
by: Qian, Keyang, et al.
Published: (2025)
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
by: Mirza, Imran, et al.
Published: (2025)
by: Mirza, Imran, et al.
Published: (2025)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
by: Dong, Zhichen, et al.
Published: (2024)
by: Dong, Zhichen, et al.
Published: (2024)
Unmasking the Canvas: A Dynamic Benchmark for Image Generation Jailbreaking and LLM Content Safety
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
by: Nair, Variath Madhupal Gautham, et al.
Published: (2025)
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
by: Potham, Ram
Published: (2025)
by: Potham, Ram
Published: (2025)
Trust in AI: Progress, Challenges, and Future Directions
by: Afroogh, Saleh, et al.
Published: (2024)
by: Afroogh, Saleh, et al.
Published: (2024)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
by: Cai, Yunna, et al.
Published: (2025)
by: Cai, Yunna, et al.
Published: (2025)
An Evaluation of Cultural Value Alignment in LLM
by: Sukiennik, Nicholas, et al.
Published: (2025)
by: Sukiennik, Nicholas, et al.
Published: (2025)
Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications
by: Goh, Jia Yi, et al.
Published: (2025)
by: Goh, Jia Yi, et al.
Published: (2025)
The CitizenQuery Benchmark: A Novel Dataset and Evaluation Pipeline for Measuring LLM Performance in Citizen Query Tasks
by: Majithia, Neil, et al.
Published: (2026)
by: Majithia, Neil, et al.
Published: (2026)
Agree to Disagree? A Meta-Evaluation of LLM Misgendering
by: Subramonian, Arjun, et al.
Published: (2025)
by: Subramonian, Arjun, et al.
Published: (2025)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Protecting Africa's Future: Cybersecurity Strategies for Child Safety, Learning, and Skill Acquisition in Tanzania
by: Gilliard, Ezekia, et al.
Published: (2024)
by: Gilliard, Ezekia, et al.
Published: (2024)
From Hallucination to Scheming: A Unified Taxonomy and Benchmark Analysis for LLM Deception
by: Shi, Jerick, et al.
Published: (2026)
by: Shi, Jerick, et al.
Published: (2026)
LLM Safety Alignment is Divergence Estimation in Disguise
by: Haldar, Rajdeep, et al.
Published: (2025)
by: Haldar, Rajdeep, et al.
Published: (2025)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026)
by: Jung, Minji, et al.
Published: (2026)
CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety
by: An, Heajun, et al.
Published: (2026)
by: An, Heajun, et al.
Published: (2026)
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
by: Judd, Nick, et al.
Published: (2025)
by: Judd, Nick, et al.
Published: (2025)
Beyond the Rubric: Cultural Misalignment in LLM Benchmarks for Sexual and Reproductive Health
by: Dey, Sumon Kanti, et al.
Published: (2025)
by: Dey, Sumon Kanti, et al.
Published: (2025)
Similar Items
-
LLMs and Childhood Safety: Identifying Risks and Proposing a Protection Framework for Safe Child-LLM Interaction
by: Jiao, Junfeng, et al.
Published: (2025) -
Evaluating LLM Safety Across Child Development Stages: A Simulated Agent Approach
by: Murali, Abhejay, et al.
Published: (2025) -
LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models
by: Jiao, Junfeng, et al.
Published: (2025) -
LLM Harms: A Taxonomy and Discussion
by: Chen, Kevin, et al.
Published: (2025) -
AI Empathy Erodes Cognitive Autonomy in Younger Users
by: Jiao, Junfeng, et al.
Published: (2026)