Gespeichert in:
| Hauptverfasser: | Wang, Junqi, Zhang, Chunhui, Li, Jiapeng, Ma, Yuxi, Niu, Lixing, Han, Jiaheng, Peng, Yujia, Zhu, Yixin, Fan, Lifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2405.11841 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
R^3-VQA: "Read the Room" by Video Social Reasoning
von: Niu, Lixing, et al.
Veröffentlicht: (2025)
von: Niu, Lixing, et al.
Veröffentlicht: (2025)
Grounding Before Generalizing: How AI Differs from Humans in Causal Transfer
von: Xiang, Liangru, et al.
Veröffentlicht: (2026)
von: Xiang, Liangru, et al.
Veröffentlicht: (2026)
A simulation-heuristics dual-process model for intuitive physics
von: Li, Shiqian, et al.
Veröffentlicht: (2025)
von: Li, Shiqian, et al.
Veröffentlicht: (2025)
Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?
von: Lu, Yi-Long, et al.
Veröffentlicht: (2025)
von: Lu, Yi-Long, et al.
Veröffentlicht: (2025)
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
von: Xu, Hainiu, et al.
Veröffentlicht: (2024)
von: Xu, Hainiu, et al.
Veröffentlicht: (2024)
Automatic Cognitive Task Generation for In-Situ Evaluation of Embodied Agents
von: He, Xinyi, et al.
Veröffentlicht: (2026)
von: He, Xinyi, et al.
Veröffentlicht: (2026)
Word Embeddings Track Social Group Changes Across 70 Years in China
von: Ma, Yuxi, et al.
Veröffentlicht: (2025)
von: Ma, Yuxi, et al.
Veröffentlicht: (2025)
Brain in a Vat: On Missing Pieces Towards Artificial General Intelligence in Large Language Models
von: Ma, Yuxi, et al.
Veröffentlicht: (2023)
von: Ma, Yuxi, et al.
Veröffentlicht: (2023)
The Cognitive Capabilities of Generative AI: A Comparative Analysis with Human Benchmarks
von: Galatzer-Levy, Isaac R., et al.
Veröffentlicht: (2024)
von: Galatzer-Levy, Isaac R., et al.
Veröffentlicht: (2024)
A Comparative Study on Automatic Coding of Medical Letters with Explainability
von: Glen, Jamie, et al.
Veröffentlicht: (2024)
von: Glen, Jamie, et al.
Veröffentlicht: (2024)
Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenliang, et al.
Veröffentlicht: (2025)
Probing and Inducing Combinational Creativity in Vision-Language Models
von: Peng, Yongqian, et al.
Veröffentlicht: (2025)
von: Peng, Yongqian, et al.
Veröffentlicht: (2025)
Nadine: An LLM-driven Intelligent Social Robot with Affective Capabilities and Human-like Memory
von: Kang, Hangyeol, et al.
Veröffentlicht: (2024)
von: Kang, Hangyeol, et al.
Veröffentlicht: (2024)
Prioritizing High-Consequence Biological Capabilities in Evaluations of Artificial Intelligence Models
von: Pannu, Jaspreet, et al.
Veröffentlicht: (2024)
von: Pannu, Jaspreet, et al.
Veröffentlicht: (2024)
How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study
von: Zhu, Shuqi, et al.
Veröffentlicht: (2026)
von: Zhu, Shuqi, et al.
Veröffentlicht: (2026)
Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
von: Sun, Yuxi, et al.
Veröffentlicht: (2025)
von: Sun, Yuxi, et al.
Veröffentlicht: (2025)
INSIGHTBUDDY-AI: Medication Extraction and Entity Linking using Large Language Models and Ensemble Learning
von: Romero, Pablo, et al.
Veröffentlicht: (2024)
von: Romero, Pablo, et al.
Veröffentlicht: (2024)
Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
von: Han, Jindong, et al.
Veröffentlicht: (2025)
von: Han, Jindong, et al.
Veröffentlicht: (2025)
Constrained Human-AI Cooperation: An Inclusive Embodied Social Intelligence Challenge
von: Du, Weihua, et al.
Veröffentlicht: (2024)
von: Du, Weihua, et al.
Veröffentlicht: (2024)
DebugBench: Evaluating Debugging Capability of Large Language Models
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game
von: Li, Lixing
Veröffentlicht: (2026)
von: Li, Lixing
Veröffentlicht: (2026)
Human and AI Perceptual Differences in Image Classification Errors
von: Liu, Minghao, et al.
Veröffentlicht: (2023)
von: Liu, Minghao, et al.
Veröffentlicht: (2023)
LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
von: Goel, Hitesh, et al.
Veröffentlicht: (2025)
von: Goel, Hitesh, et al.
Veröffentlicht: (2025)
DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
von: Zhuang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhuang, Tianyi, et al.
Veröffentlicht: (2025)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
von: Ma, Shuai, et al.
Veröffentlicht: (2024)
von: Ma, Shuai, et al.
Veröffentlicht: (2024)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
A Conceptual Framework for AI Capability Evaluations
von: Carro, María Victoria, et al.
Veröffentlicht: (2025)
von: Carro, María Victoria, et al.
Veröffentlicht: (2025)
Replicating Human Social Perception in Generative AI: Evaluating the Valence-Dominance Model
von: Gurkan, Necdet, et al.
Veröffentlicht: (2025)
von: Gurkan, Necdet, et al.
Veröffentlicht: (2025)
MSDiagnosis: A Benchmark for Evaluating Large Language Models in Multi-Step Clinical Diagnosis
von: Hou, Ruihui, et al.
Veröffentlicht: (2024)
von: Hou, Ruihui, et al.
Veröffentlicht: (2024)
SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific Applications
von: Herron, Emily, et al.
Veröffentlicht: (2025)
von: Herron, Emily, et al.
Veröffentlicht: (2025)
Proposing and solving olympiad geometry with guided tree search
von: Zhang, Chi, et al.
Veröffentlicht: (2024)
von: Zhang, Chi, et al.
Veröffentlicht: (2024)
LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2025)
An Insight into Security Code Review with LLMs: Capabilities, Obstacles, and Influential Factors
von: Yu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Yu, Jiaxin, et al.
Veröffentlicht: (2024)
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
Exploring the Capability of ChatGPT to Reproduce Human Labels for Social Computing Tasks (Extended Version)
von: Zhu, Yiming, et al.
Veröffentlicht: (2024)
von: Zhu, Yiming, et al.
Veröffentlicht: (2024)
BugBlitz-AI: An Intelligent QA Assistant
von: Yao, Yi, et al.
Veröffentlicht: (2024)
von: Yao, Yi, et al.
Veröffentlicht: (2024)
Evaluating LLMs Capabilities Towards Understanding Social Dynamics
von: Tahir, Anique, et al.
Veröffentlicht: (2024)
von: Tahir, Anique, et al.
Veröffentlicht: (2024)
Open-World Evaluations for Measuring Frontier AI Capabilities
von: Kapoor, Sayash, et al.
Veröffentlicht: (2026)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2026)
Sheet as Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding
von: Lei, Yiming, et al.
Veröffentlicht: (2026)
von: Lei, Yiming, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
R^3-VQA: "Read the Room" by Video Social Reasoning
von: Niu, Lixing, et al.
Veröffentlicht: (2025) -
Grounding Before Generalizing: How AI Differs from Humans in Causal Transfer
von: Xiang, Liangru, et al.
Veröffentlicht: (2026) -
A simulation-heuristics dual-process model for intuitive physics
von: Li, Shiqian, et al.
Veröffentlicht: (2025) -
Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?
von: Lu, Yi-Long, et al.
Veröffentlicht: (2025) -
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
von: Xu, Hainiu, et al.
Veröffentlicht: (2024)