Evaluating Large Language Models on the Frame and Symbol Grounding Problems: A Zero-shot Benchmark
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Oka, Shoko |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
ZeroDL: Zero-shot Distribution Learning for Text Clustering via Large Language Models
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024)
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024)
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
von: Li, Zekun, et al.
Veröffentlicht: (2024)
von: Li, Zekun, et al.
Veröffentlicht: (2024)
ZeFaV: Boosting Large Language Models for Zero-shot Fact Verification
von: Luu, Son T., et al.
Veröffentlicht: (2024)
von: Luu, Son T., et al.
Veröffentlicht: (2024)
Text2World: Benchmarking Large Language Models for Symbolic World Model Generation
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
von: Hu, Mengkang, et al.
Veröffentlicht: (2025)
LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
von: Shibata, Takumi, et al.
Veröffentlicht: (2025)
von: Shibata, Takumi, et al.
Veröffentlicht: (2025)
DeFrame: Debiasing Large Language Models Against Framing Effects
von: Lim, Kahee, et al.
Veröffentlicht: (2026)
von: Lim, Kahee, et al.
Veröffentlicht: (2026)
A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical Tasks
von: Labrak, Yanis, et al.
Veröffentlicht: (2023)
von: Labrak, Yanis, et al.
Veröffentlicht: (2023)
How Large Language Models Need Symbolism
von: Deng, Xiaotie, et al.
Veröffentlicht: (2025)
von: Deng, Xiaotie, et al.
Veröffentlicht: (2025)
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2023)
Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models
von: Xu, Fangzhi, et al.
Veröffentlicht: (2023)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2023)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
von: Colle, Vincenzo, et al.
Veröffentlicht: (2025)
von: Colle, Vincenzo, et al.
Veröffentlicht: (2025)
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
TurkBench: A Benchmark for Evaluating Turkish Large Language Models
von: Toraman, Çağrı, et al.
Veröffentlicht: (2026)
von: Toraman, Çağrı, et al.
Veröffentlicht: (2026)
EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
von: Bae, Suyoung, et al.
Veröffentlicht: (2025)
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
von: Xu, Xinnuo, et al.
Veröffentlicht: (2025)
von: Xu, Xinnuo, et al.
Veröffentlicht: (2025)
Benchmarking Open-Source Large Language Models for Persian in Zero-Shot and Few-Shot Learning
von: Cherakhloo, Mahdi, et al.
Veröffentlicht: (2025)
von: Cherakhloo, Mahdi, et al.
Veröffentlicht: (2025)
PCEval: A Benchmark for Evaluating Physical Computing Capabilities of Large Language Models
von: Song, Inpyo, et al.
Veröffentlicht: (2025)
von: Song, Inpyo, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models to Extract Information on Substance Use Disorder Severity from Clinical Notes: A Zero-shot Learning Approach
von: Mahbub, Maria, et al.
Veröffentlicht: (2024)
von: Mahbub, Maria, et al.
Veröffentlicht: (2024)
CodeApex: A Bilingual Programming Evaluation Benchmark for Large Language Models
von: Fu, Lingyue, et al.
Veröffentlicht: (2023)
von: Fu, Lingyue, et al.
Veröffentlicht: (2023)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset
von: Zhu, Jie, et al.
Veröffentlicht: (2024)
von: Zhu, Jie, et al.
Veröffentlicht: (2024)
Mathify: Evaluating Large Language Models on Mathematical Problem Solving Tasks
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
von: Tang, Weizhi, et al.
Veröffentlicht: (2024)
Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution
von: Jin, Qiao, et al.
Veröffentlicht: (2026)
von: Jin, Qiao, et al.
Veröffentlicht: (2026)
In Context Learning and Reasoning for Symbolic Regression with Large Language Models
von: Sharlin, Samiha, et al.
Veröffentlicht: (2024)
von: Sharlin, Samiha, et al.
Veröffentlicht: (2024)
The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models
von: Singh, Abhinav Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Abhinav Kumar, et al.
Veröffentlicht: (2026)
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology
von: Zhou, Chengfeng, et al.
Veröffentlicht: (2025)
von: Zhou, Chengfeng, et al.
Veröffentlicht: (2025)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
von: Guo, Qianhong, et al.
Veröffentlicht: (2025)
von: Guo, Qianhong, et al.
Veröffentlicht: (2025)
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics
von: Somasekharan, Nithin, et al.
Veröffentlicht: (2025)
von: Somasekharan, Nithin, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models for Financial Reasoning: A CFA-Based Benchmark Study
von: Yao, Xuan, et al.
Veröffentlicht: (2025)
von: Yao, Xuan, et al.
Veröffentlicht: (2025)
MoZIP: A Multilingual Benchmark to Evaluate Large Language Models in Intellectual Property
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
von: Ni, Shiwen, et al.
Veröffentlicht: (2024)
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
von: Jahan, Israt, et al.
Veröffentlicht: (2023)
von: Jahan, Israt, et al.
Veröffentlicht: (2023)
WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models
von: Ning, Kangyun, et al.
Veröffentlicht: (2024)
von: Ning, Kangyun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
von: Pombal, José, et al.
Veröffentlicht: (2025) -
ZeroDL: Zero-shot Distribution Learning for Text Clustering via Large Language Models
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2024) -
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2023) -
Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models
von: Liu, Qiang, et al.
Veröffentlicht: (2025) -
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
von: Li, Zekun, et al.
Veröffentlicht: (2024)