Enregistré dans:
| Auteurs principaux: | Chen, Zhuang, Cao, Yaru, Bi, Guanqun, Wu, Jincenzi, Zhou, Jinfeng, Xiao, Xiyao, Chen, Si, Wang, Hongning, Huang, Minlie |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2506.16756 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ToMBench: Benchmarking Theory of Mind in Large Language Models
par: Chen, Zhuang, et autres
Publié: (2024)
par: Chen, Zhuang, et autres
Publié: (2024)
PsychePass: Calibrating LLM Therapeutic Competence via Trajectory-Anchored Tournaments
par: Chen, Zhuang, et autres
Publié: (2026)
par: Chen, Zhuang, et autres
Publié: (2026)
Think Socially via Cognitive Reasoning
par: Zhou, Jinfeng, et autres
Publié: (2025)
par: Zhou, Jinfeng, et autres
Publié: (2025)
Unveiling the Landscape of Clinical Depression Assessment: From Behavioral Signatures to Psychiatric Reasoning
par: Chen, Zhuang, et autres
Publié: (2025)
par: Chen, Zhuang, et autres
Publié: (2025)
CharacterBench: Benchmarking Character Customization of Large Language Models
par: Zhou, Jinfeng, et autres
Publié: (2024)
par: Zhou, Jinfeng, et autres
Publié: (2024)
SS-GEN: A Social Story Generation Framework with Large Language Models
par: Feng, Yi, et autres
Publié: (2024)
par: Feng, Yi, et autres
Publié: (2024)
COKE: A Cognitive Knowledge Graph for Machine Theory of Mind
par: Wu, Jincenzi, et autres
Publié: (2023)
par: Wu, Jincenzi, et autres
Publié: (2023)
SocialEval: Evaluating Social Intelligence of Large Language Models
par: Zhou, Jinfeng, et autres
Publié: (2025)
par: Zhou, Jinfeng, et autres
Publié: (2025)
MAGI: Multi-Agent Guided Interview for Psychiatric Assessment
par: Bi, Guanqun, et autres
Publié: (2025)
par: Bi, Guanqun, et autres
Publié: (2025)
Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite Feedback
par: Zhu, Shijing, et autres
Publié: (2025)
par: Zhu, Shijing, et autres
Publié: (2025)
Data Selection via Optimal Control for Language Models
par: Gu, Yuxian, et autres
Publié: (2024)
par: Gu, Yuxian, et autres
Publié: (2024)
Crisp: Cognitive Restructuring of Negative Thoughts through Multi-turn Supportive Dialogues
par: Zhou, Jinfeng, et autres
Publié: (2025)
par: Zhou, Jinfeng, et autres
Publié: (2025)
Language Model Decoding as Direct Metrics Optimization
par: Ji, Haozhe, et autres
Publié: (2023)
par: Ji, Haozhe, et autres
Publié: (2023)
Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning
par: Wen, Jiaxin, et autres
Publié: (2024)
par: Wen, Jiaxin, et autres
Publié: (2024)
Reframe Your Life Story: Interactive Narrative Therapist and Innovative Moment Assessment with Large Language Models
par: Feng, Yi, et autres
Publié: (2025)
par: Feng, Yi, et autres
Publié: (2025)
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators
par: Wen, Bosi, et autres
Publié: (2025)
par: Wen, Bosi, et autres
Publié: (2025)
Sentipolis: Emotion-Aware Agents for Social Simulations
par: Fu, Chiyuan, et autres
Publié: (2026)
par: Fu, Chiyuan, et autres
Publié: (2026)
Towards Optimal Learning of Language Models
par: Gu, Yuxian, et autres
Publié: (2024)
par: Gu, Yuxian, et autres
Publié: (2024)
AMOR: A Recipe for Building Adaptable Modular Knowledge Agents Through Process Feedback
par: Guan, Jian, et autres
Publié: (2024)
par: Guan, Jian, et autres
Publié: (2024)
Learning Task Decomposition to Assist Humans in Competitive Programming
par: Wen, Jiaxin, et autres
Publié: (2024)
par: Wen, Jiaxin, et autres
Publié: (2024)
Mitigating Strategy Preference Bias in Emotional Support Conversation via Uncertainty Estimations
par: Zhou, Yougen, et autres
Publié: (2025)
par: Zhou, Yougen, et autres
Publié: (2025)
Towards Efficient Exact Optimization of Language Model Alignment
par: Ji, Haozhe, et autres
Publié: (2024)
par: Ji, Haozhe, et autres
Publié: (2024)
Grounding LLMs in Scientific Discovery via Embodied Actions
par: Zhang, Bo, et autres
Publié: (2026)
par: Zhang, Bo, et autres
Publié: (2026)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
par: Zhang, Zhexin, et autres
Publié: (2024)
par: Zhang, Zhexin, et autres
Publié: (2024)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
par: Zhang, Zhexin, et autres
Publié: (2023)
par: Zhang, Zhexin, et autres
Publié: (2023)
SocialBench: Sociality Evaluation of Role-Playing Conversational Agents
par: Chen, Hongzhan, et autres
Publié: (2024)
par: Chen, Hongzhan, et autres
Publié: (2024)
A Group Fairness Lens for Large Language Models
par: Bi, Guanqun, et autres
Publié: (2023)
par: Bi, Guanqun, et autres
Publié: (2023)
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models
par: Cheng, Jiale, et autres
Publié: (2024)
par: Cheng, Jiale, et autres
Publié: (2024)
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
par: Wen, Bosi, et autres
Publié: (2025)
par: Wen, Bosi, et autres
Publié: (2025)
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
par: Cheng, Jiale, et autres
Publié: (2023)
par: Cheng, Jiale, et autres
Publié: (2023)
Social-R1: Towards Human-like Social Reasoning in LLMs
par: Wu, Jincenzi, et autres
Publié: (2026)
par: Wu, Jincenzi, et autres
Publié: (2026)
VicSim: Enhancing Victim Simulation with Emotional and Linguistic Fidelity
par: Li, Yerong, et autres
Publié: (2025)
par: Li, Yerong, et autres
Publié: (2025)
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
par: Zhang, Zhexin, et autres
Publié: (2025)
par: Zhang, Zhexin, et autres
Publié: (2025)
ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models
par: Zhao, Haiquan, et autres
Publié: (2024)
par: Zhao, Haiquan, et autres
Publié: (2024)
Towards Exception Safety Code Generation with Intermediate Representation Agents Framework
par: Zhang, Xuanming, et autres
Publié: (2024)
par: Zhang, Xuanming, et autres
Publié: (2024)
CARE: Cognitive-reasoning Augmented Reinforcement for Emotional Support Conversation
par: Zhu, Jie, et autres
Publié: (2025)
par: Zhu, Jie, et autres
Publié: (2025)
Benchmarking Complex Instruction-Following with Multiple Constraints Composition
par: Wen, Bosi, et autres
Publié: (2024)
par: Wen, Bosi, et autres
Publié: (2024)
EmoBench: Evaluating the Emotional Intelligence of Large Language Models
par: Sabour, Sahand, et autres
Publié: (2024)
par: Sabour, Sahand, et autres
Publié: (2024)
RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
par: Feng, Andrew Zhuoer, et autres
Publié: (2026)
par: Feng, Andrew Zhuoer, et autres
Publié: (2026)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
par: Cui, Shiyao, et autres
Publié: (2025)
par: Cui, Shiyao, et autres
Publié: (2025)
Documents similaires
-
ToMBench: Benchmarking Theory of Mind in Large Language Models
par: Chen, Zhuang, et autres
Publié: (2024) -
PsychePass: Calibrating LLM Therapeutic Competence via Trajectory-Anchored Tournaments
par: Chen, Zhuang, et autres
Publié: (2026) -
Think Socially via Cognitive Reasoning
par: Zhou, Jinfeng, et autres
Publié: (2025) -
Unveiling the Landscape of Clinical Depression Assessment: From Behavioral Signatures to Psychiatric Reasoning
par: Chen, Zhuang, et autres
Publié: (2025) -
CharacterBench: Benchmarking Character Customization of Large Language Models
par: Zhou, Jinfeng, et autres
Publié: (2024)