Your Co-Workers Matter: Evaluating Collaborative Capabilities of Language Models in Blocks World
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Guande, Zhao, Chen, Silva, Claudio, He, He |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
Automated Novelty Evaluation of Academic Paper: A Collaborative Approach Integrating Human and Large Language Model Knowledge
von: Wu, Wenqing, et al.
Veröffentlicht: (2025)
von: Wu, Wenqing, et al.
Veröffentlicht: (2025)
From Control to Foresight: Simulation as a New Paradigm for Human-Agent Collaboration
von: He, Gaole, et al.
Veröffentlicht: (2026)
von: He, Gaole, et al.
Veröffentlicht: (2026)
StressPrompt: Does Stress Impact Large Language Models and Human Performance Similarly?
von: Shen, Guobin, et al.
Veröffentlicht: (2024)
von: Shen, Guobin, et al.
Veröffentlicht: (2024)
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
von: Zhao, Runcong, et al.
Veröffentlicht: (2025)
von: Zhao, Runcong, et al.
Veröffentlicht: (2025)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
WundtGPT: Shaping Large Language Models To Be An Empathetic, Proactive Psychologist
von: Ren, Chenyu, et al.
Veröffentlicht: (2024)
von: Ren, Chenyu, et al.
Veröffentlicht: (2024)
Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools
von: Hao, Yilun, et al.
Veröffentlicht: (2024)
von: Hao, Yilun, et al.
Veröffentlicht: (2024)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
von: Liu, Hongtao, et al.
Veröffentlicht: (2025)
An Evaluation of Estimative Uncertainty in Large Language Models
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhisheng, et al.
Veröffentlicht: (2024)
Evaluating the Prompt Steerability of Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
Collaborative Evaluation of Deepfake Text with Deliberation-Enhancing Dialogue Systems
von: Lee, Jooyoung, et al.
Veröffentlicht: (2025)
von: Lee, Jooyoung, et al.
Veröffentlicht: (2025)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
von: Sun, Chongyan, et al.
Veröffentlicht: (2024)
von: Sun, Chongyan, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models in Analysing Classroom Dialogue
von: Long, Yun, et al.
Veröffentlicht: (2024)
von: Long, Yun, et al.
Veröffentlicht: (2024)
POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models
von: He, Jianben, et al.
Veröffentlicht: (2024)
von: He, Jianben, et al.
Veröffentlicht: (2024)
A Scalable Framework for Evaluating Health Language Models
von: Mallinar, Neil, et al.
Veröffentlicht: (2025)
von: Mallinar, Neil, et al.
Veröffentlicht: (2025)
Direct Advantage Regression: Aligning LLMs with Online AI Reward
von: He, Li, et al.
Veröffentlicht: (2025)
von: He, Li, et al.
Veröffentlicht: (2025)
Evalet: Evaluating Large Language Models through Functional Fragmentation
von: Kim, Tae Soo, et al.
Veröffentlicht: (2025)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2025)
Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
von: Yao, Bingsheng, et al.
Veröffentlicht: (2025)
von: Yao, Bingsheng, et al.
Veröffentlicht: (2025)
Autograding Mathematical Induction Proofs with Natural Language Processing
von: Zhao, Chenyan, et al.
Veröffentlicht: (2024)
von: Zhao, Chenyan, et al.
Veröffentlicht: (2024)
Prompts Matter: Comparing ML/GAI Approaches for Generating Inductive Qualitative Coding Results
von: Chen, John, et al.
Veröffentlicht: (2024)
von: Chen, John, et al.
Veröffentlicht: (2024)
Stories of Your Life as Others: A Round-Trip Evaluation of LLM-Generated Life Stories Conditioned on Rich Psychometric Profiles
von: Wigler, Ben, et al.
Veröffentlicht: (2026)
von: Wigler, Ben, et al.
Veröffentlicht: (2026)
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments
von: Daynauth, Roland, et al.
Veröffentlicht: (2024)
von: Daynauth, Roland, et al.
Veröffentlicht: (2024)
Large Language Model-Brained GUI Agents: A Survey
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
Helmsman of the Masses? Evaluate the Opinion Leadership of Large Language Models in the Werewolf Game
von: Du, Silin, et al.
Veröffentlicht: (2024)
von: Du, Silin, et al.
Veröffentlicht: (2024)
Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement
von: Ye, Haoran, et al.
Veröffentlicht: (2025)
von: Ye, Haoran, et al.
Veröffentlicht: (2025)
The Effectiveness of Style Vectors for Steering Large Language Models: A Human Evaluation
von: Diallo, Diaoulé, et al.
Veröffentlicht: (2026)
von: Diallo, Diaoulé, et al.
Veröffentlicht: (2026)
A Survey on Human-AI Collaboration with Large Foundation Models
von: Vats, Vanshika, et al.
Veröffentlicht: (2024)
von: Vats, Vanshika, et al.
Veröffentlicht: (2024)
Human Evaluation of Procedural Knowledge Graph Extraction from Text with Large Language Models
von: Carriero, Valentina Anita, et al.
Veröffentlicht: (2024)
von: Carriero, Valentina Anita, et al.
Veröffentlicht: (2024)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2023)
Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
von: Qian, Cheng, et al.
Veröffentlicht: (2024)
Generative Interfaces for Language Models
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2025)
When AI Writes, Whose Voice Remains? Quantifying Cultural Marker Erasure Across World English Varieties in Large Language Models
von: Navneet, Satyam Kumar, et al.
Veröffentlicht: (2026)
von: Navneet, Satyam Kumar, et al.
Veröffentlicht: (2026)
RNR: Teaching Large Language Models to Follow Roles and Rules
von: Wang, Kuan, et al.
Veröffentlicht: (2024)
von: Wang, Kuan, et al.
Veröffentlicht: (2024)
Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
von: Zhu, Shaojie, et al.
Veröffentlicht: (2023)
von: Zhu, Shaojie, et al.
Veröffentlicht: (2023)
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
von: Lee, Suhyun, et al.
Veröffentlicht: (2026)
von: Lee, Suhyun, et al.
Veröffentlicht: (2026)
Art or Artifice? Large Language Models and the False Promise of Creativity
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2023)
von: Chakrabarty, Tuhin, et al.
Veröffentlicht: (2023)
Evaluating Large Language Models' Ability Using a Psychiatric Screening Tool Based on Metaphor and Sarcasm Scenarios
von: Yakura, Hiromu
Veröffentlicht: (2023)
von: Yakura, Hiromu
Veröffentlicht: (2023)
Satori: Towards Proactive AR Assistant with Belief-Desire-Intention User Modeling
von: Li, Chenyi, et al.
Veröffentlicht: (2024)
von: Li, Chenyi, et al.
Veröffentlicht: (2024)
CoCo Matrix: Taxonomy of Cognitive Contributions in Co-writing with Intelligent Agents
von: Wan, Ruyuan, et al.
Veröffentlicht: (2024)
von: Wan, Ruyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
von: Wang, Kangyu, et al.
Veröffentlicht: (2025) -
Automated Novelty Evaluation of Academic Paper: A Collaborative Approach Integrating Human and Large Language Model Knowledge
von: Wu, Wenqing, et al.
Veröffentlicht: (2025) -
From Control to Foresight: Simulation as a New Paradigm for Human-Agent Collaboration
von: He, Gaole, et al.
Veröffentlicht: (2026) -
StressPrompt: Does Stress Impact Large Language Models and Human Performance Similarly?
von: Shen, Guobin, et al.
Veröffentlicht: (2024) -
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
von: Zhao, Runcong, et al.
Veröffentlicht: (2025)