Re-evaluating Theory of Mind evaluation in large language models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Jennifer, Sosa, Felix, Ullman, Tomer |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Cognitive models can reveal interpretable value trade-offs in language models
di: Murthy, Sonia K., et al.
Pubblicazione: (2025)
di: Murthy, Sonia K., et al.
Pubblicazione: (2025)
Shades of Zero: Distinguishing Impossibility from Inconceivability
di: Hu, Jennifer, et al.
Pubblicazione: (2025)
di: Hu, Jennifer, et al.
Pubblicazione: (2025)
TeleEval-OS: Performance evaluations of large language models for operations scheduling
di: Wang, Yanyan, et al.
Pubblicazione: (2025)
di: Wang, Yanyan, et al.
Pubblicazione: (2025)
Can open source large language models be used for tumor documentation in Germany? -- An evaluation on urological doctors' notes
di: Lenz, Stefan, et al.
Pubblicazione: (2025)
di: Lenz, Stefan, et al.
Pubblicazione: (2025)
MMToM-QA: Multimodal Theory of Mind Question Answering
di: Jin, Chuanyang, et al.
Pubblicazione: (2024)
di: Jin, Chuanyang, et al.
Pubblicazione: (2024)
One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity
di: Murthy, Sonia K., et al.
Pubblicazione: (2024)
di: Murthy, Sonia K., et al.
Pubblicazione: (2024)
Morphological evaluation of subwords vocabulary used by BETO language model
di: García-Sierra, Óscar, et al.
Pubblicazione: (2024)
di: García-Sierra, Óscar, et al.
Pubblicazione: (2024)
CVE-LLM : Automatic vulnerability evaluation in medical device industry using large language models
di: Ghosh, Rikhiya, et al.
Pubblicazione: (2024)
di: Ghosh, Rikhiya, et al.
Pubblicazione: (2024)
MindScope: Exploring cognitive biases in large language models through Multi-Agent Systems
di: Xie, Zhentao, et al.
Pubblicazione: (2024)
di: Xie, Zhentao, et al.
Pubblicazione: (2024)
Dissociating language and thought in large language models
di: Mahowald, Kyle, et al.
Pubblicazione: (2023)
di: Mahowald, Kyle, et al.
Pubblicazione: (2023)
Forking Paths in Neural Text Generation
di: Bigelow, Eric, et al.
Pubblicazione: (2024)
di: Bigelow, Eric, et al.
Pubblicazione: (2024)
Auxiliary task demands mask the capabilities of smaller language models
di: Hu, Jennifer, et al.
Pubblicazione: (2024)
di: Hu, Jennifer, et al.
Pubblicazione: (2024)
ReIFE: Re-evaluating Instruction-Following Evaluation
di: Liu, Yixin, et al.
Pubblicazione: (2024)
di: Liu, Yixin, et al.
Pubblicazione: (2024)
On the attribution of confidence to large language models
di: Keeling, Geoff, et al.
Pubblicazione: (2024)
di: Keeling, Geoff, et al.
Pubblicazione: (2024)
ChildEval: When large language models meet children's personalities
di: Luo, Yanyan, et al.
Pubblicazione: (2026)
di: Luo, Yanyan, et al.
Pubblicazione: (2026)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
di: Gao, Mingqi, et al.
Pubblicazione: (2024)
The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic
di: Długosz, Dominika Agnieszka, et al.
Pubblicazione: (2026)
di: Długosz, Dominika Agnieszka, et al.
Pubblicazione: (2026)
Quantifying non deterministic drift in large language models
di: Nicholson, Claire
Pubblicazione: (2026)
di: Nicholson, Claire
Pubblicazione: (2026)
Can large language models build causal graphs?
di: Long, Stephanie, et al.
Pubblicazione: (2023)
di: Long, Stephanie, et al.
Pubblicazione: (2023)
Multi-round jailbreak attack on large language models
di: Zhou, Yihua, et al.
Pubblicazione: (2024)
di: Zhou, Yihua, et al.
Pubblicazione: (2024)
Response: Emergent analogical reasoning in large language models
di: Hodel, Damian, et al.
Pubblicazione: (2023)
di: Hodel, Damian, et al.
Pubblicazione: (2023)
The 20 questions game to distinguish large language models
di: Richardeau, Gurvan, et al.
Pubblicazione: (2024)
di: Richardeau, Gurvan, et al.
Pubblicazione: (2024)
Representation in large language models
di: Yetman, Cameron
Pubblicazione: (2025)
di: Yetman, Cameron
Pubblicazione: (2025)
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning
di: Wagner, Eitan, et al.
Pubblicazione: (2024)
di: Wagner, Eitan, et al.
Pubblicazione: (2024)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
di: Janiak, Denis, et al.
Pubblicazione: (2025)
di: Janiak, Denis, et al.
Pubblicazione: (2025)
Long-form factuality in large language models
di: Wei, Jerry, et al.
Pubblicazione: (2024)
di: Wei, Jerry, et al.
Pubblicazione: (2024)
Failure of contextual invariance in large language models
di: Kumar, Sagar, et al.
Pubblicazione: (2026)
di: Kumar, Sagar, et al.
Pubblicazione: (2026)
Chain or tree? Re-evaluating complex reasoning from the perspective of a matrix of thought
di: Tang, Fengxiao, et al.
Pubblicazione: (2025)
di: Tang, Fengxiao, et al.
Pubblicazione: (2025)
FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models
di: Shamsfard, Mehrnoush, et al.
Pubblicazione: (2025)
di: Shamsfard, Mehrnoush, et al.
Pubblicazione: (2025)
A survey of textual cyber abuse detection using cutting-edge language models and large language models
di: Diaz-Garcia, Jose A., et al.
Pubblicazione: (2025)
di: Diaz-Garcia, Jose A., et al.
Pubblicazione: (2025)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
di: Albadarneh, Israa A., et al.
Pubblicazione: (2025)
di: Albadarneh, Israa A., et al.
Pubblicazione: (2025)
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
di: Padlewski, Piotr, et al.
Pubblicazione: (2024)
di: Padlewski, Piotr, et al.
Pubblicazione: (2024)
Streamlining evidence based clinical recommendations with large language models
di: Li, Dubai, et al.
Pubblicazione: (2025)
di: Li, Dubai, et al.
Pubblicazione: (2025)
Strong and weak alignment of large language models with human values
di: Khamassi, Mehdi, et al.
Pubblicazione: (2024)
di: Khamassi, Mehdi, et al.
Pubblicazione: (2024)
Correcting misinformation on social media with a large language model
di: Zhou, Xinyi, et al.
Pubblicazione: (2024)
di: Zhou, Xinyi, et al.
Pubblicazione: (2024)
A review on the use of large language models as virtual tutors
di: García-Méndez, Silvia, et al.
Pubblicazione: (2024)
di: García-Méndez, Silvia, et al.
Pubblicazione: (2024)
Evaluating large language models in medical applications: a survey
di: Chen, Xiaolan, et al.
Pubblicazione: (2024)
di: Chen, Xiaolan, et al.
Pubblicazione: (2024)
Disentangling generalization and memorization in large language models using chess
di: Pleiss, Leonard S., et al.
Pubblicazione: (2026)
di: Pleiss, Leonard S., et al.
Pubblicazione: (2026)
MathDivide: Improved mathematical reasoning by large language models
di: Srivastava, Saksham Sahai, et al.
Pubblicazione: (2024)
di: Srivastava, Saksham Sahai, et al.
Pubblicazione: (2024)
An evaluation of LLMs and Google Translate for translation of selected Indian languages via sentiment and semantic analyses
di: Chandra, Rohitash, et al.
Pubblicazione: (2025)
di: Chandra, Rohitash, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Cognitive models can reveal interpretable value trade-offs in language models
di: Murthy, Sonia K., et al.
Pubblicazione: (2025) -
Shades of Zero: Distinguishing Impossibility from Inconceivability
di: Hu, Jennifer, et al.
Pubblicazione: (2025) -
TeleEval-OS: Performance evaluations of large language models for operations scheduling
di: Wang, Yanyan, et al.
Pubblicazione: (2025) -
Can open source large language models be used for tumor documentation in Germany? -- An evaluation on urological doctors' notes
di: Lenz, Stefan, et al.
Pubblicazione: (2025) -
MMToM-QA: Multimodal Theory of Mind Question Answering
di: Jin, Chuanyang, et al.
Pubblicazione: (2024)