PCEval: A Benchmark for Evaluating Physical Computing Capabilities of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Inpyo, Jeon, Eunji, Lee, Jangwon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
von: Li, Ce, et al.
Veröffentlicht: (2025)
von: Li, Ce, et al.
Veröffentlicht: (2025)
A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist
von: Jeon, Sohyeon, et al.
Veröffentlicht: (2025)
von: Jeon, Sohyeon, et al.
Veröffentlicht: (2025)
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
von: Xu, Hainiu, et al.
Veröffentlicht: (2024)
von: Xu, Hainiu, et al.
Veröffentlicht: (2024)
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
von: Pauli, Amalie Brogaard, et al.
Veröffentlicht: (2024)
von: Pauli, Amalie Brogaard, et al.
Veröffentlicht: (2024)
ElectriQ: A Benchmark for Assessing the Response Capability of Large Language Models in Power Marketing
von: Wang, Jinzhi, et al.
Veröffentlicht: (2025)
von: Wang, Jinzhi, et al.
Veröffentlicht: (2025)
Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
von: Kim, Dahyun, et al.
Veröffentlicht: (2024)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics
von: Somasekharan, Nithin, et al.
Veröffentlicht: (2025)
von: Somasekharan, Nithin, et al.
Veröffentlicht: (2025)
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
Instance-Aligned Captions for Explainable Video Anomaly Detection
von: Song, Inpyo, et al.
Veröffentlicht: (2026)
von: Song, Inpyo, et al.
Veröffentlicht: (2026)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
von: Bohnet, Bernd, et al.
Veröffentlicht: (2024)
SQLBench: A Comprehensive Evaluation for Text-to-SQL Capabilities of Large Language Models
von: Zhang, Bin, et al.
Veröffentlicht: (2024)
von: Zhang, Bin, et al.
Veröffentlicht: (2024)
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
von: Jang, Kyochul, et al.
Veröffentlicht: (2025)
von: Jang, Kyochul, et al.
Veröffentlicht: (2025)
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
von: Holtermann, Carolin, et al.
Veröffentlicht: (2024)
von: Holtermann, Carolin, et al.
Veröffentlicht: (2024)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
MedFact: Benchmarking the Fact-Checking Capabilities of Large Language Models on Chinese Medical Texts
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
von: Zbeeb, Mohammad, et al.
Veröffentlicht: (2025)
von: Zbeeb, Mohammad, et al.
Veröffentlicht: (2025)
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2023)
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models
von: Lee, Jaeho, et al.
Veröffentlicht: (2025)
von: Lee, Jaeho, et al.
Veröffentlicht: (2025)
Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games
von: Malik, Saumya
Veröffentlicht: (2024)
von: Malik, Saumya
Veröffentlicht: (2024)
TurkBench: A Benchmark for Evaluating Turkish Large Language Models
von: Toraman, Çağrı, et al.
Veröffentlicht: (2026)
von: Toraman, Çağrı, et al.
Veröffentlicht: (2026)
EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design
von: Neehal, Nafis, et al.
Veröffentlicht: (2024)
von: Neehal, Nafis, et al.
Veröffentlicht: (2024)
DebugBench: Evaluating Debugging Capability of Large Language Models
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
von: Wang, Wentian, et al.
Veröffentlicht: (2024)
von: Wang, Wentian, et al.
Veröffentlicht: (2024)
GraphInstruct: Empowering Large Language Models with Graph Understanding and Reasoning Capability
von: Luo, Zihan, et al.
Veröffentlicht: (2024)
von: Luo, Zihan, et al.
Veröffentlicht: (2024)
Real-time Traffic Accident Anticipation with Feature Reuse
von: Song, Inpyo, et al.
Veröffentlicht: (2025)
von: Song, Inpyo, et al.
Veröffentlicht: (2025)
Bounding-Box Trajectories Matter for Video Anomaly Detection
von: Song, Inpyo, et al.
Veröffentlicht: (2026)
von: Song, Inpyo, et al.
Veröffentlicht: (2026)
CodeApex: A Bilingual Programming Evaluation Benchmark for Large Language Models
von: Fu, Lingyue, et al.
Veröffentlicht: (2023)
von: Fu, Lingyue, et al.
Veröffentlicht: (2023)
FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs
von: Lee, Shinbok, et al.
Veröffentlicht: (2024)
von: Lee, Shinbok, et al.
Veröffentlicht: (2024)
Evaluating Interventional Reasoning Capabilities of Large Language Models
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
Reasoning Capabilities of Large Language Models on Dynamic Tasks
von: Wong, Annie, et al.
Veröffentlicht: (2025)
von: Wong, Annie, et al.
Veröffentlicht: (2025)
Do Large Language Models Know What They Are Capable Of?
von: Barkan, Casey O., et al.
Veröffentlicht: (2025)
von: Barkan, Casey O., et al.
Veröffentlicht: (2025)
Exploring Large Language Models for Multimodal Sentiment Analysis: Challenges, Benchmarks, and Future Directions
von: Song, Shezheng
Veröffentlicht: (2024)
von: Song, Shezheng
Veröffentlicht: (2024)
On the Use of Large Language Models to Generate Capability Ontologies
von: da Silva, Luis Miguel Vieira, et al.
Veröffentlicht: (2024)
von: da Silva, Luis Miguel Vieira, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
von: Li, Ce, et al.
Veröffentlicht: (2025) -
A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist
von: Jeon, Sohyeon, et al.
Veröffentlicht: (2025) -
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024) -
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
von: Xu, Hainiu, et al.
Veröffentlicht: (2024) -
Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
von: Pauli, Amalie Brogaard, et al.
Veröffentlicht: (2024)