How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO
Fuente:
arXiv
Saved in:
| Main Authors: | Ng, Man Tik, Tse, Hui Tung, Huang, Jen-tse, Li, Jingjing, Wang, Wenxuan, Lyu, Michael R. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
by: Huang, Jen-tse, et al.
Published: (2024)
by: Huang, Jen-tse, et al.
Published: (2024)
ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
by: Li, Shuqing, et al.
Published: (2025)
by: Li, Shuqing, et al.
Published: (2025)
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
by: Huang, Jen-tse, et al.
Published: (2023)
by: Huang, Jen-tse, et al.
Published: (2023)
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
by: Du, Yongkang, et al.
Published: (2025)
by: Du, Yongkang, et al.
Published: (2025)
On the Shortcut Learning in Multilingual Neural Machine Translation
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
by: Lam, Man Ho, et al.
Published: (2025)
by: Lam, Man Ho, et al.
Published: (2025)
Revisiting the Reliability of Psychological Scales on Large Language Models
by: Huang, Jen-tse, et al.
Published: (2023)
by: Huang, Jen-tse, et al.
Published: (2023)
On the Failure of Latent State Persistence in Large Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
AI Sees Your Location, But With A Bias Toward The Wealthy World
by: Huang, Jingyuan, et al.
Published: (2025)
by: Huang, Jingyuan, et al.
Published: (2025)
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench
by: Huang, Jen-tse, et al.
Published: (2023)
by: Huang, Jen-tse, et al.
Published: (2023)
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
by: Liu, Xiaoyuan, et al.
Published: (2024)
by: Liu, Xiaoyuan, et al.
Published: (2024)
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
by: Yuan, Youliang, et al.
Published: (2023)
by: Yuan, Youliang, et al.
Published: (2023)
DROGO: Default Representation Objective via Graph Optimization in Reinforcement Learning
by: Tse, Hon Tik, et al.
Published: (2026)
by: Tse, Hon Tik, et al.
Published: (2026)
Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews
by: Wang, Xintao, et al.
Published: (2023)
by: Wang, Xintao, et al.
Published: (2023)
All Languages Matter: On the Multilingual Safety of Large Language Models
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
The Achilles' Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities
by: Qin, Zixuan, et al.
Published: (2025)
by: Qin, Zixuan, et al.
Published: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024)
by: Yuan, Youliang, et al.
Published: (2024)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
by: Wang, Xintao, et al.
Published: (2025)
by: Wang, Xintao, et al.
Published: (2025)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
by: Huang, Jen-tse, et al.
Published: (2025)
by: Huang, Jen-tse, et al.
Published: (2025)
Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
by: Yuan, Youliang, et al.
Published: (2025)
by: Yuan, Youliang, et al.
Published: (2025)
InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game Context
by: Liu, Ziyi, et al.
Published: (2024)
by: Liu, Ziyi, et al.
Published: (2024)
Reward-Aware Proto-Representations in Reinforcement Learning
by: Tse, Hon Tik, et al.
Published: (2025)
by: Tse, Hon Tik, et al.
Published: (2025)
The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
by: Zhou, Jiaxu, et al.
Published: (2025)
by: Zhou, Jiaxu, et al.
Published: (2025)
RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
by: Wang, Zekun Moore, et al.
Published: (2023)
by: Wang, Zekun Moore, et al.
Published: (2023)
New Job, New Gender? Measuring the Social Bias in Image Generation Models
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation
by: Huang, Jen-tse, et al.
Published: (2026)
by: Huang, Jen-tse, et al.
Published: (2026)
How Easily Can AI Chatbots Spread Misinformation in Audiology and Otolaryngology?
by: W. Wiktor Jedrzejczak, et al.
Published: (2026)
by: W. Wiktor Jedrzejczak, et al.
Published: (2026)
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
by: Huang, Jen-tse, et al.
Published: (2026)
by: Huang, Jen-tse, et al.
Published: (2026)
A Framework for Evaluating Appropriateness, Trustworthiness, and Safety in Mental Wellness AI Chatbots
by: Chen, Lucia, et al.
Published: (2024)
by: Chen, Lucia, et al.
Published: (2024)
Exploring the Impact of Anthropomorphism in Role-Playing AI Chatbots on Media Dependency: A Case Study of Xuanhe AI
by: Yu, Qiufang, et al.
Published: (2024)
by: Yu, Qiufang, et al.
Published: (2024)
How Well Can a Long Sequence Model Model Long Sequences? Comparing Architechtural Inductive Biases on Long-Context Abilities
by: Huang, Jerry
Published: (2024)
by: Huang, Jerry
Published: (2024)
Orca: Enhancing Role-Playing Abilities of Large Language Models by Integrating Personality Traits
by: Huang, Yuxuan
Published: (2024)
by: Huang, Yuxuan
Published: (2024)
Can Models Help Us Create Better Models? Evaluating LLMs as Data Scientists
by: Pietruszka, Michał, et al.
Published: (2024)
by: Pietruszka, Michał, et al.
Published: (2024)
Reasoning Does Not Necessarily Improve Role-Playing Ability
by: Feng, Xiachong, et al.
Published: (2025)
by: Feng, Xiachong, et al.
Published: (2025)
Similar Items
-
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
by: Huang, Jen-tse, et al.
Published: (2024) -
ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
by: Li, Shuqing, et al.
Published: (2025) -
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
by: Huang, Jen-tse, et al.
Published: (2023) -
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
by: Wan, Yuxuan, et al.
Published: (2024) -
Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases
by: Huang, Jen-tse, et al.
Published: (2025)