Cross-Platform Evaluation of Large Language Model Safety in Pediatric Consultations: Evolution of Adversarial Robustness and the Scale Paradox
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Zolfaghari, Vahideh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PediatricAnxietyBench: Evaluating Large Language Model Safety Under Parental Anxiety and Pressure in Pediatric Consultations
von: Zolfaghari, Vahideh
Veröffentlicht: (2025)
von: Zolfaghari, Vahideh
Veröffentlicht: (2025)
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
von: Zolfaghari, Vahideh
Veröffentlicht: (2026)
von: Zolfaghari, Vahideh
Veröffentlicht: (2026)
On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models
von: Yang, April, et al.
Veröffentlicht: (2024)
von: Yang, April, et al.
Veröffentlicht: (2024)
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
von: Diao, Muxi, et al.
Veröffentlicht: (2024)
von: Diao, Muxi, et al.
Veröffentlicht: (2024)
LongSafety: Evaluating Long-Context Safety of Large Language Models
von: Lu, Yida, et al.
Veröffentlicht: (2025)
von: Lu, Yida, et al.
Veröffentlicht: (2025)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
KGPA: Robustness Evaluation for Large Language Models via Cross-Domain Knowledge Graphs
von: Pei, Aihua, et al.
Veröffentlicht: (2024)
von: Pei, Aihua, et al.
Veröffentlicht: (2024)
Evaluating Psychological Safety of Large Language Models
von: Li, Xingxuan, et al.
Veröffentlicht: (2022)
von: Li, Xingxuan, et al.
Veröffentlicht: (2022)
Adversarial Reinforcement Learning for Large Language Model Agent Safety
von: Wang, Zizhao, et al.
Veröffentlicht: (2025)
von: Wang, Zizhao, et al.
Veröffentlicht: (2025)
Evaluating the Retrieval Robustness of Large Language Models
von: Cao, Shuyang, et al.
Veröffentlicht: (2025)
von: Cao, Shuyang, et al.
Veröffentlicht: (2025)
A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models
von: Feier, Andrei Marian, et al.
Veröffentlicht: (2026)
von: Feier, Andrei Marian, et al.
Veröffentlicht: (2026)
Towards Safety Evaluations of Theory of Mind in Large Language Models
von: Aoshima, Tatsuhiro, et al.
Veröffentlicht: (2025)
von: Aoshima, Tatsuhiro, et al.
Veröffentlicht: (2025)
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
von: Galisai, Marcello, et al.
Veröffentlicht: (2026)
von: Galisai, Marcello, et al.
Veröffentlicht: (2026)
Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
von: de Curtò, J., et al.
Veröffentlicht: (2025)
von: de Curtò, J., et al.
Veröffentlicht: (2025)
Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
von: McCoy, Liam G., et al.
Veröffentlicht: (2025)
von: McCoy, Liam G., et al.
Veröffentlicht: (2025)
The Rosetta Paradox: Domain-Specific Performance Inversions in Large Language Models
von: Jha, Basab, et al.
Veröffentlicht: (2024)
von: Jha, Basab, et al.
Veröffentlicht: (2024)
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
von: Walkowiak, Paweł, et al.
Veröffentlicht: (2025)
von: Walkowiak, Paweł, et al.
Veröffentlicht: (2025)
Knowledge-Graph Based RAG System Evaluation Framework
von: Dong, Sicheng, et al.
Veröffentlicht: (2025)
von: Dong, Sicheng, et al.
Veröffentlicht: (2025)
Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
von: Park, Sungwon, et al.
Veröffentlicht: (2024)
von: Park, Sungwon, et al.
Veröffentlicht: (2024)
SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
von: Torres-Fonseca, Josue, et al.
Veröffentlicht: (2026)
von: Torres-Fonseca, Josue, et al.
Veröffentlicht: (2026)
EMRModel: A Large Language Model for Extracting Medical Consultation Dialogues into Structured Medical Records
von: Zhao, Shuguang, et al.
Veröffentlicht: (2025)
von: Zhao, Shuguang, et al.
Veröffentlicht: (2025)
A Role-specific Guided Large Language Model for Ophthalmic Consultation Based on Stylistic Differentiation
von: Fu, Laiyi, et al.
Veröffentlicht: (2024)
von: Fu, Laiyi, et al.
Veröffentlicht: (2024)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
von: Röttger, Paul, et al.
Veröffentlicht: (2024)
von: Röttger, Paul, et al.
Veröffentlicht: (2024)
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
von: Gupta, Prannaya, et al.
Veröffentlicht: (2024)
von: Gupta, Prannaya, et al.
Veröffentlicht: (2024)
Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
von: Chen, Zijian, et al.
Veröffentlicht: (2025)
von: Chen, Zijian, et al.
Veröffentlicht: (2025)
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
von: Zong, Xuanjun, et al.
Veröffentlicht: (2025)
von: Zong, Xuanjun, et al.
Veröffentlicht: (2025)
More Women, Same Stereotypes: Unpacking the Gender Bias Paradox in Large Language Models
von: Chen, Evan, et al.
Veröffentlicht: (2025)
von: Chen, Evan, et al.
Veröffentlicht: (2025)
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
von: Villa, Danielle, et al.
Veröffentlicht: (2025)
von: Villa, Danielle, et al.
Veröffentlicht: (2025)
Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation
von: Cantini, Riccardo, et al.
Veröffentlicht: (2024)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2024)
A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks
von: Nguyen, Hieu Minh "Jord"
Veröffentlicht: (2025)
von: Nguyen, Hieu Minh "Jord"
Veröffentlicht: (2025)
CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference
von: Yu, Erxin, et al.
Veröffentlicht: (2024)
von: Yu, Erxin, et al.
Veröffentlicht: (2024)
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models for Generalization and Robustness via Data Compression
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
von: Li, Yucheng, et al.
Veröffentlicht: (2024)
Performance Evaluation of Lightweight Open-source Large Language Models in Pediatric Consultations: A Comparative Analysis
von: Wei, Qiuhong, et al.
Veröffentlicht: (2024)
von: Wei, Qiuhong, et al.
Veröffentlicht: (2024)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
von: Cantini, Riccardo, et al.
Veröffentlicht: (2025)
Quantifying Label-Induced Bias in Large Language Model Self- and Cross-Evaluations
von: Saraf, Muskan, et al.
Veröffentlicht: (2025)
von: Saraf, Muskan, et al.
Veröffentlicht: (2025)
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
von: Jackson, Declan, et al.
Veröffentlicht: (2025)
von: Jackson, Declan, et al.
Veröffentlicht: (2025)
MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues
von: Liu, Zheyuan, et al.
Veröffentlicht: (2026)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2026)
All Languages Matter: On the Multilingual Safety of Large Language Models
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
PediatricAnxietyBench: Evaluating Large Language Model Safety Under Parental Anxiety and Pressure in Pediatric Consultations
von: Zolfaghari, Vahideh
Veröffentlicht: (2025) -
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
von: Zolfaghari, Vahideh
Veröffentlicht: (2026) -
On Adversarial Robustness and Out-of-Distribution Robustness of Large Language Models
von: Yang, April, et al.
Veröffentlicht: (2024) -
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
von: Diao, Muxi, et al.
Veröffentlicht: (2024) -
LongSafety: Evaluating Long-Context Safety of Large Language Models
von: Lu, Yida, et al.
Veröffentlicht: (2025)