MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Hoyun, Kang, Migyeong, Shin, Jisu, Kim, Jihyun, Park, Chanbi, Yoo, Hangyeol, An, Jihyun, Oh, Alice, Han, Jinyoung, Lim, KyungTae |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TREX: Tokenizer Regression for Optimal Data Mixture
by: Won, Inho, et al.
Published: (2026)
by: Won, Inho, et al.
Published: (2026)
SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification
by: Kang, Migyeong, et al.
Published: (2026)
by: Kang, Migyeong, et al.
Published: (2026)
Before and After ChatGPT: Revisiting AI-Based Dialogue Systems for Emotional Support
by: Lee, Daeun, et al.
Published: (2026)
by: Lee, Daeun, et al.
Published: (2026)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring
by: Song, Jayoung, et al.
Published: (2025)
by: Song, Jayoung, et al.
Published: (2025)
Korean Named Entity Recognition Based on Language-Specific Features
by: Chen, Yige, et al.
Published: (2023)
by: Chen, Yige, et al.
Published: (2023)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
by: Lim, Hyeonseok, et al.
Published: (2024)
by: Lim, Hyeonseok, et al.
Published: (2024)
Do matching skill sets matter? Exploring the relationship between motivation, satisfaction, and role identification in the context of older adults' volunteering
by: Jihyun Park
Published: (2024)
by: Jihyun Park
Published: (2024)
Enhancing Korean Dependency Parsing with Morphosyntactic Features
by: Park, Jungyeul, et al.
Published: (2025)
by: Park, Jungyeul, et al.
Published: (2025)
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation
by: Lee, Huije, et al.
Published: (2026)
by: Lee, Huije, et al.
Published: (2026)
Does Rationale Quality Matter? Enhancing Mental Disorder Detection via Selective Reasoning Distillation
by: Song, Hoyun, et al.
Published: (2025)
by: Song, Hoyun, et al.
Published: (2025)
Learning Constituent Headedness
by: Qi, Zeyao, et al.
Published: (2026)
by: Qi, Zeyao, et al.
Published: (2026)
Enhanced Conditional Generation of Double Perovskite by Knowledge-Guided Language Model Feedback
by: Lee, Inhyo, et al.
Published: (2025)
by: Lee, Inhyo, et al.
Published: (2025)
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
by: Hwang, Bokwang, et al.
Published: (2025)
by: Hwang, Bokwang, et al.
Published: (2025)
Log Heston Model for Monthly Average VIX
by: Park, Jihyun, et al.
Published: (2024)
by: Park, Jihyun, et al.
Published: (2024)
Zero-Coupon Treasury Rates and Returns using the Volatility Index
by: Park, Jihyun, et al.
Published: (2024)
by: Park, Jihyun, et al.
Published: (2024)
The VIX as Stochastic Volatility for Corporate Bonds
by: Park, Jihyun, et al.
Published: (2024)
by: Park, Jihyun, et al.
Published: (2024)
NOE analysis using dual injection DNP‐NMR for studies of solvent–solute interactions at low concentrations
by: Jihyun Kim
Published: (2024)
by: Jihyun Kim
Published: (2024)
Recent Findings from the Telescope Array Experiment
by: Kim, Jihyun
Published: (2026)
by: Kim, Jihyun
Published: (2026)
Developing an instrument to measure Korean pre‐service teachers' understanding of language as an epistemic tool in mathematics education
by: Jihyun Hwang
Published: (2024)
by: Jihyun Hwang
Published: (2024)
KORMo: Korean Open Reasoning Model for Everyone
by: Kim, Minjun, et al.
Published: (2025)
by: Kim, Minjun, et al.
Published: (2025)
MENTOR: A Reinforcement Learning Framework for Enabling Tool Use in Small Models via Teacher-Optimized Rewards
by: Choi, ChangSu, et al.
Published: (2025)
by: Choi, ChangSu, et al.
Published: (2025)
K-UD: Revising Korean Universal Dependencies Guidelines
by: Kim, Kyuwon, et al.
Published: (2024)
by: Kim, Kyuwon, et al.
Published: (2024)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach
by: Ko, Changgeon, et al.
Published: (2024)
by: Ko, Changgeon, et al.
Published: (2024)
Topical dexamethasone decelerates epithelial migration on the canine tympanic membrane
by: Jihyun Kim, et al.
Published: (2024)
by: Jihyun Kim, et al.
Published: (2024)
Asynchronous Sharpness-Aware Minimization For Fast and Accurate Deep Learning
by: Jo, Junhyuk, et al.
Published: (2025)
by: Jo, Junhyuk, et al.
Published: (2025)
MPMAvatar: Learning 3D Gaussian Avatars with Accurate and Robust Physics-Based Dynamics
by: Lee, Changmin, et al.
Published: (2025)
by: Lee, Changmin, et al.
Published: (2025)
MedRegion-CT: Region-Focused Multimodal LLM for Comprehensive 3D CT Report Generation
by: Kyung, Sunggu, et al.
Published: (2025)
by: Kyung, Sunggu, et al.
Published: (2025)
Associations between dietary patterns, gut microbiome diversity, and itch severity in preschool aged children with atopic dermatitis: A cross‐sectional study
by: Jisu Park, et al.
Published: (2025)
by: Jisu Park, et al.
Published: (2025)
Multi‐Dimensional Physically Unclonable Functions: Optoelectronic Variation‐Induced Multi‐Key Generation from Small Molecule PN Heterostructures
by: Jihyun Shin, et al.
Published: (2024)
by: Jihyun Shin, et al.
Published: (2024)
MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports
by: Kyung, Sunggu, et al.
Published: (2025)
by: Kyung, Sunggu, et al.
Published: (2025)
ELO: Efficient Layer-Specific Optimization for Continual Pretraining of Multilingual LLMs
by: Yoo, HanGyeol, et al.
Published: (2026)
by: Yoo, HanGyeol, et al.
Published: (2026)
Real world prescription of beta‐blockers in patients with asthma
by: Jihyun Lee, et al.
Published: (2024)
by: Jihyun Lee, et al.
Published: (2024)
Efficacy of Mindfulness‐Based Interventions for Reducing Cancer‐Related Fatigue: A Systematic Review and Meta‐Analysis
by: Jihyun Lee, et al.
Published: (2026)
by: Jihyun Lee, et al.
Published: (2026)
An unexpected property of $\mathbf{g}$-vectors for rank 3 mutation-cyclic quivers
by: Lee, Jihyun, et al.
Published: (2024)
by: Lee, Jihyun, et al.
Published: (2024)
Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
by: Moon, Jihyun, et al.
Published: (2025)
by: Moon, Jihyun, et al.
Published: (2025)
Ultrathin Gallium Oxide as Both Surface Passivation Layer with Conductive Filament Contacts and Alternative Gate Dielectric for 2D MOSFETs
by: Sanghyun Moon, et al.
Published: (2025)
by: Sanghyun Moon, et al.
Published: (2025)
Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM Collectives
by: Ko, Changgeon, et al.
Published: (2026)
by: Ko, Changgeon, et al.
Published: (2026)
Similar Items
-
TREX: Tokenizer Regression for Optimal Data Mixture
by: Won, Inho, et al.
Published: (2026) -
SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification
by: Kang, Migyeong, et al.
Published: (2026) -
Before and After ChatGPT: Revisiting AI-Based Dialogue Systems for Emotional Support
by: Lee, Daeun, et al.
Published: (2026) -
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025) -
Enriching the Korean Learner Corpus with Multi-reference Annotations and Rubric-Based Scoring
by: Song, Jayoung, et al.
Published: (2025)