Evaluating LLMs with Multiple Problems at once
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhengxiang, Kodner, Jordan, Rambow, Owen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English Writing
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2025)
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2025)
Clustering Document Parts: Detecting and Characterizing Influence Campaigns from Documents
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2024)
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2024)
Evaluating Neural Language Models as Cognitive Models of Language Acquisition
von: Martínez, Héctor Javier Vázquez, et al.
Veröffentlicht: (2023)
von: Martínez, Héctor Javier Vázquez, et al.
Veröffentlicht: (2023)
Examining Gender and Power on Wikipedia Through Face and Politeness
von: Soubki, Adil, et al.
Veröffentlicht: (2024)
von: Soubki, Adil, et al.
Veröffentlicht: (2024)
LVLMs and Humans Ground Differently in Referential Communication
von: Zeng, Peter, et al.
Veröffentlicht: (2026)
von: Zeng, Peter, et al.
Veröffentlicht: (2026)
Zero-Shot Belief: A Hard Problem for LLMs
von: Murzaku, John, et al.
Veröffentlicht: (2025)
von: Murzaku, John, et al.
Veröffentlicht: (2025)
Training LLMs to Recognize Hedges in Spontaneous Narratives
von: Paige, Amie J., et al.
Veröffentlicht: (2024)
von: Paige, Amie J., et al.
Veröffentlicht: (2024)
Measuring Iterative Temporal Reasoning with Time Puzzles
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2026)
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2026)
Synthetic Audio Helps for Cognitive State Tasks
von: Soubki, Adil, et al.
Veröffentlicht: (2025)
von: Soubki, Adil, et al.
Veröffentlicht: (2025)
NormSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-Fly
von: Fung, Yi R., et al.
Veröffentlicht: (2022)
von: Fung, Yi R., et al.
Veröffentlicht: (2022)
Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2025)
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2025)
OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs
von: Murzaku, John, et al.
Veröffentlicht: (2025)
von: Murzaku, John, et al.
Veröffentlicht: (2025)
LVLMs are Bad at Overhearing Human Referential Communication
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2025)
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2025)
Evaluating and Enhancing LLMs for Multi-turn Text-to-SQL with Multiple Question Types
von: Guo, Ziming, et al.
Veröffentlicht: (2024)
von: Guo, Ziming, et al.
Veröffentlicht: (2024)
Optimizing Length Compression in Large Reasoning Models
von: Cheng, Zhengxiang, et al.
Veröffentlicht: (2025)
von: Cheng, Zhengxiang, et al.
Veröffentlicht: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
von: Henkel, Owen, et al.
Veröffentlicht: (2024)
von: Henkel, Owen, et al.
Veröffentlicht: (2024)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
von: Li, Hanyu, et al.
Veröffentlicht: (2025)
DePT: Decomposed Prompt Tuning for Parameter-Efficient Fine-tuning
von: Shi, Zhengxiang, et al.
Veröffentlicht: (2023)
von: Shi, Zhengxiang, et al.
Veröffentlicht: (2023)
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
Intention and Face in Dialog
von: Soubki, Adil, et al.
Veröffentlicht: (2024)
von: Soubki, Adil, et al.
Veröffentlicht: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
Pretrained LLMs Learn Multiple Types of Uncertainty
von: Cohen, Roi, et al.
Veröffentlicht: (2025)
von: Cohen, Roi, et al.
Veröffentlicht: (2025)
Visual Reasoning Benchmark: Evaluating Multimodal LLMs on Classroom-Authentic Visual Problems from Primary Education
von: Huti, Mohamed, et al.
Veröffentlicht: (2026)
von: Huti, Mohamed, et al.
Veröffentlicht: (2026)
Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems
von: Ye, Junyi, et al.
Veröffentlicht: (2024)
von: Ye, Junyi, et al.
Veröffentlicht: (2024)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
von: Xiao, Yuxin, et al.
Veröffentlicht: (2024)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2024)
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
DHP Benchmark: Are LLMs Good NLG Evaluators?
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset
von: Henkel, Owen, et al.
Veröffentlicht: (2023)
von: Henkel, Owen, et al.
Veröffentlicht: (2023)
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
von: Yang, Chang, et al.
Veröffentlicht: (2025)
von: Yang, Chang, et al.
Veröffentlicht: (2025)
Wait, that's not an option: LLMs Robustness with Incorrect Multiple-Choice Options
von: Góral, Gracjan, et al.
Veröffentlicht: (2024)
von: Góral, Gracjan, et al.
Veröffentlicht: (2024)
Cascaded Self-Evaluation Augmented Training for Lightweight Multimodal LLMs
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning
von: Dinucu-Jianu, David, et al.
Veröffentlicht: (2025)
von: Dinucu-Jianu, David, et al.
Veröffentlicht: (2025)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
von: Murugadoss, Bhuvanashree, et al.
Veröffentlicht: (2024)
von: Murugadoss, Bhuvanashree, et al.
Veröffentlicht: (2024)
LLMs versus the Halting Problem: Characterizing Program Termination Reasoning
von: Sultan, Oren, et al.
Veröffentlicht: (2026)
von: Sultan, Oren, et al.
Veröffentlicht: (2026)
Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems
von: Peng, Dengyun, et al.
Veröffentlicht: (2025)
von: Peng, Dengyun, et al.
Veröffentlicht: (2025)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
von: Huang, Jen-tse, et al.
Veröffentlicht: (2024)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English Writing
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2025) -
Clustering Document Parts: Detecting and Characterizing Influence Campaigns from Documents
von: Wang, Zhengxiang, et al.
Veröffentlicht: (2024) -
Evaluating Neural Language Models as Cognitive Models of Language Acquisition
von: Martínez, Héctor Javier Vázquez, et al.
Veröffentlicht: (2023) -
Examining Gender and Power on Wikipedia Through Face and Politeness
von: Soubki, Adil, et al.
Veröffentlicht: (2024) -
LVLMs and Humans Ground Differently in Referential Communication
von: Zeng, Peter, et al.
Veröffentlicht: (2026)