Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines
Fuente:
arXiv
Salvato in:
| Autori principali: | Deng, Guifeng, Rao, Shuyin, Lin, Tianyu, Dai, Anlu, Wang, Pan, Xie, Junyi, Song, Haidong, Zhao, Ke, Xu, Dongwu, Cheng, Zhengdong, Li, Tao, Jiang, Haiteng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model
di: Deng, Guifeng, et al.
Pubblicazione: (2026)
di: Deng, Guifeng, et al.
Pubblicazione: (2026)
An Exploratory Deep Learning Approach for Predicting Subsequent Suicidal Acts in Chinese Psychological Support Hotlines
di: Song, Changwei, et al.
Pubblicazione: (2024)
di: Song, Changwei, et al.
Pubblicazione: (2024)
Deep Learning-Based Feature Fusion for Emotion Analysis and Suicide Risk Differentiation in Chinese Psychological Support Hotlines
di: Wang, Han, et al.
Pubblicazione: (2025)
di: Wang, Han, et al.
Pubblicazione: (2025)
Deep Learning and Large Language Models for Audio and Text Analysis in Predicting Suicidal Acts in Chinese Psychological Support Hotlines
di: Chen, Yining, et al.
Pubblicazione: (2024)
di: Chen, Yining, et al.
Pubblicazione: (2024)
Fine-grained Speech Sentiment Analysis in Chinese Psychological Support Hotlines Based on Large-scale Pre-trained Model
di: Chen, Zhonglong, et al.
Pubblicazione: (2024)
di: Chen, Zhonglong, et al.
Pubblicazione: (2024)
Learning Nonlinear Systems In-Context: From Synthetic Data to Real-World Motor Control
di: Jian, Tong, et al.
Pubblicazione: (2026)
di: Jian, Tong, et al.
Pubblicazione: (2026)
“Digital Friend” or “It”? Conceptualizations of LLM ‐Powered Chatbots in National Sexual Assault and Domestic Violence Crisis Hotlines
di: Nikki M. Wise
Pubblicazione: (2025)
di: Nikki M. Wise
Pubblicazione: (2025)
PsyScam: A Benchmark for Psychological Techniques in Real-World Scams
di: Ma, Shang, et al.
Pubblicazione: (2025)
di: Ma, Shang, et al.
Pubblicazione: (2025)
Add a Donor Hotline for Giving Society Members
Pubblicazione: (2025)
Pubblicazione: (2025)
Evaluation of the Effect of Dietary Manganese on the Intestinal Digestive Function, Antioxidant Response, and Muscle Quality in Coho Salmon
di: Dongwu Liu, et al.
Pubblicazione: (2024)
di: Dongwu Liu, et al.
Pubblicazione: (2024)
Research on the Fitting Method for P‐S‐N Curves With Extremely Small Sample Experiment Data: Improved Backwards Statistical Inference Method
di: Tong Mu, et al.
Pubblicazione: (2025)
di: Tong Mu, et al.
Pubblicazione: (2025)
The Effects of Organizational Dependence and Hotline Administration on Whistleblowing Accounting Fraud
di: Eva Zedlacher, et al.
Pubblicazione: (2025)
di: Eva Zedlacher, et al.
Pubblicazione: (2025)
BioMaze: Benchmarking and Enhancing Large Language Models for Biological Pathway Reasoning
di: Zhao, Haiteng, et al.
Pubblicazione: (2025)
di: Zhao, Haiteng, et al.
Pubblicazione: (2025)
WeatherBench: A Real-World Benchmark Dataset for All-in-One Adverse Weather Image Restoration
di: Guan, Qiyuan, et al.
Pubblicazione: (2025)
di: Guan, Qiyuan, et al.
Pubblicazione: (2025)
Mary Kay Ash Foundation renews partnership with National Domestic Violence Hotline.
Pubblicazione: (2025)
Pubblicazione: (2025)
High-Contrast Interferometric Imaging of Single-Molecule Dynamics on Optical Fibers
di: Li, Guifeng, et al.
Pubblicazione: (2025)
di: Li, Guifeng, et al.
Pubblicazione: (2025)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
di: Yang, Shuo, et al.
Pubblicazione: (2025)
di: Yang, Shuo, et al.
Pubblicazione: (2025)
Approximate Borderline Sampling using Granular-Ball for Classification Tasks
di: Xie, Qin, et al.
Pubblicazione: (2025)
di: Xie, Qin, et al.
Pubblicazione: (2025)
La Autoridad Personal en el Sistema Familiar: Adaptación y Validación a la Población Mexicana
di: Shuyin Durán Torres
Pubblicazione: (2012)
di: Shuyin Durán Torres
Pubblicazione: (2012)
Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
di: Mai, Wuyuao, et al.
Pubblicazione: (2025)
di: Mai, Wuyuao, et al.
Pubblicazione: (2025)
Beyond Query-Level Comparison: Fine-Grained Reinforcement Learning for Text-to-SQL with Automated Interpretable Critiques
di: Wang, Guifeng, et al.
Pubblicazione: (2025)
di: Wang, Guifeng, et al.
Pubblicazione: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
di: Ding, Shuangrui, et al.
Pubblicazione: (2026)
di: Ding, Shuangrui, et al.
Pubblicazione: (2026)
Generative Adversarial Network on Motion-Blur Image Restoration
di: Li, Zhengdong
Pubblicazione: (2024)
di: Li, Zhengdong
Pubblicazione: (2024)
A Survey on Blockchain-based Supply Chain Finance with Progress and Future directions
di: Luo, Zhengdong
Pubblicazione: (2024)
di: Luo, Zhengdong
Pubblicazione: (2024)
An Overview of Machine Learning-Driven Resource Allocation in IoT Networks
di: Li, Zhengdong
Pubblicazione: (2024)
di: Li, Zhengdong
Pubblicazione: (2024)
Towards Evaluation for Real-World LLM Unlearning
di: Miao, Ke, et al.
Pubblicazione: (2025)
di: Miao, Ke, et al.
Pubblicazione: (2025)
Granular-ball Representation Learning for Deep CNN on Learning with Label Noise
di: Dai, Dawei, et al.
Pubblicazione: (2024)
di: Dai, Dawei, et al.
Pubblicazione: (2024)
Generalizable Sleep Staging via Multi-Level Domain Alignment
di: Wang, Jiquan, et al.
Pubblicazione: (2023)
di: Wang, Jiquan, et al.
Pubblicazione: (2023)
EEGAgent: A Unified Framework for Automated EEG Analysis Using Large Language Models
di: Zhao, Sha, et al.
Pubblicazione: (2025)
di: Zhao, Sha, et al.
Pubblicazione: (2025)
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
di: Song, Zhiheng, et al.
Pubblicazione: (2026)
di: Song, Zhiheng, et al.
Pubblicazione: (2026)
SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities
di: Yoash, Noga Ben, et al.
Pubblicazione: (2025)
di: Yoash, Noga Ben, et al.
Pubblicazione: (2025)
MLN-net: A multi-source medical image segmentation method for clustered microcalcifications using multiple layer normalization
di: Wang, Ke, et al.
Pubblicazione: (2023)
di: Wang, Ke, et al.
Pubblicazione: (2023)
RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
di: Chen, Yuhao, et al.
Pubblicazione: (2026)
di: Chen, Yuhao, et al.
Pubblicazione: (2026)
GBSVM: Granular-ball Support Vector Machine
di: Xia, Shuyin, et al.
Pubblicazione: (2022)
di: Xia, Shuyin, et al.
Pubblicazione: (2022)
Instruction-Based Molecular Graph Generation with Unified Text-Graph Diffusion Model
di: Xiang, Yuran, et al.
Pubblicazione: (2024)
di: Xiang, Yuran, et al.
Pubblicazione: (2024)
CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation
di: Sawarni, Ayush, et al.
Pubblicazione: (2026)
di: Sawarni, Ayush, et al.
Pubblicazione: (2026)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
di: Wang, Yanlin, et al.
Pubblicazione: (2026)
di: Wang, Yanlin, et al.
Pubblicazione: (2026)
CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments
di: Yu, Yi, et al.
Pubblicazione: (2026)
di: Yu, Yi, et al.
Pubblicazione: (2026)
PolyReal: A Benchmark for Real-World Polymer Science Workflows
di: Liu, Wanhao, et al.
Pubblicazione: (2026)
di: Liu, Wanhao, et al.
Pubblicazione: (2026)
Toward Real-World Chinese Psychological Support Dialogues: CPsDD Dataset and a Co-Evolving Multi-Agent System
di: Shi, Yuanchen, et al.
Pubblicazione: (2025)
di: Shi, Yuanchen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model
di: Deng, Guifeng, et al.
Pubblicazione: (2026) -
An Exploratory Deep Learning Approach for Predicting Subsequent Suicidal Acts in Chinese Psychological Support Hotlines
di: Song, Changwei, et al.
Pubblicazione: (2024) -
Deep Learning-Based Feature Fusion for Emotion Analysis and Suicide Risk Differentiation in Chinese Psychological Support Hotlines
di: Wang, Han, et al.
Pubblicazione: (2025) -
Deep Learning and Large Language Models for Audio and Text Analysis in Predicting Suicidal Acts in Chinese Psychological Support Hotlines
di: Chen, Yining, et al.
Pubblicazione: (2024) -
Fine-grained Speech Sentiment Analysis in Chinese Psychological Support Hotlines Based on Large-scale Pre-trained Model
di: Chen, Zhonglong, et al.
Pubblicazione: (2024)