A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task
Fuente:
arXiv
Salvato in:
| Autori principali: | Fujisaki, Yuya, Takagi, Shiro, Asoh, Hideki, Kumagai, Wataru |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
di: Hashimoto, Wataru, et al.
Pubblicazione: (2024)
di: Hashimoto, Wataru, et al.
Pubblicazione: (2024)
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
di: Park, Jin Hyun, et al.
Pubblicazione: (2025)
di: Park, Jin Hyun, et al.
Pubblicazione: (2025)
IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
di: Mohanty, Shrestha, et al.
Pubblicazione: (2024)
di: Mohanty, Shrestha, et al.
Pubblicazione: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
di: Ouyang, Jialin
Pubblicazione: (2025)
di: Ouyang, Jialin
Pubblicazione: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
di: He, Yuanqin, et al.
Pubblicazione: (2024)
di: He, Yuanqin, et al.
Pubblicazione: (2024)
Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models
di: Fuchi, Masane, et al.
Pubblicazione: (2025)
di: Fuchi, Masane, et al.
Pubblicazione: (2025)
Extraction of Research Objectives, Machine Learning Model Names, and Dataset Names from Academic Papers and Analysis of Their Interrelationships Using LLM and Network Analysis
di: Nishio, S., et al.
Pubblicazione: (2024)
di: Nishio, S., et al.
Pubblicazione: (2024)
Evaluating Fine-Tuned LLM Model For Medical Transcription With Small Low-Resource Languages Validated Dataset
di: Chowdhury, Mohammed Nowshad Ruhani, et al.
Pubblicazione: (2026)
di: Chowdhury, Mohammed Nowshad Ruhani, et al.
Pubblicazione: (2026)
HindSight: Evaluating LLM-Generated Research Ideas via Future Impact
di: Jiang, Bo
Pubblicazione: (2026)
di: Jiang, Bo
Pubblicazione: (2026)
M-QUEST -- Meme Question-Understanding Evaluation on Semantics and Toxicity
di: De Giorgis, Stefano, et al.
Pubblicazione: (2026)
di: De Giorgis, Stefano, et al.
Pubblicazione: (2026)
Graphical Reasoning: LLM-based Semi-Open Relation Extraction
di: Tao, Yicheng, et al.
Pubblicazione: (2024)
di: Tao, Yicheng, et al.
Pubblicazione: (2024)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
di: Wang, Zhanliang, et al.
Pubblicazione: (2026)
di: Wang, Zhanliang, et al.
Pubblicazione: (2026)
There Are No Silly Questions: Evaluation of Offline LLM Capabilities from a Turkish Perspective
di: Yilmaz, Edibe, et al.
Pubblicazione: (2026)
di: Yilmaz, Edibe, et al.
Pubblicazione: (2026)
Training on the Test Task Confounds Evaluation and Emergence
di: Dominguez-Olmedo, Ricardo, et al.
Pubblicazione: (2024)
di: Dominguez-Olmedo, Ricardo, et al.
Pubblicazione: (2024)
ChaI-TeA: A Benchmark for Evaluating Autocompletion of Interactions with LLM-based Chatbots
di: Goren, Shani, et al.
Pubblicazione: (2024)
di: Goren, Shani, et al.
Pubblicazione: (2024)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
di: Roytburg, Dani, et al.
Pubblicazione: (2026)
di: Roytburg, Dani, et al.
Pubblicazione: (2026)
Enhancing LLM Evaluations: The Garbling Trick
di: Bradley, William F.
Pubblicazione: (2024)
di: Bradley, William F.
Pubblicazione: (2024)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Evaluating Language-Model Agents on Realistic Autonomous Tasks
di: Kinniment, Megan, et al.
Pubblicazione: (2023)
di: Kinniment, Megan, et al.
Pubblicazione: (2023)
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation
di: Gong, Albert, et al.
Pubblicazione: (2025)
di: Gong, Albert, et al.
Pubblicazione: (2025)
A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian
di: Rogoz, Ana-Cristina, et al.
Pubblicazione: (2025)
di: Rogoz, Ana-Cristina, et al.
Pubblicazione: (2025)
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2024)
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2024)
Towards Multilingual LLM Evaluation for European Languages
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
di: Patel, Bhrij, et al.
Pubblicazione: (2024)
di: Patel, Bhrij, et al.
Pubblicazione: (2024)
Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
di: Agnimo, Yedidia, et al.
Pubblicazione: (2026)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
Mathematical Derivation Graphs: A Relation Extraction Task in STEM Manuscripts
di: Prasad, Vishesh, et al.
Pubblicazione: (2024)
di: Prasad, Vishesh, et al.
Pubblicazione: (2024)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
di: Gupta, Vipul, et al.
Pubblicazione: (2024)
di: Gupta, Vipul, et al.
Pubblicazione: (2024)
MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 Languages
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
Rhetorical Questions in LLM Representations: A Linear Probing Study
di: Yao, Louie Hong, et al.
Pubblicazione: (2026)
di: Yao, Louie Hong, et al.
Pubblicazione: (2026)
CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks
di: Feng, Jie, et al.
Pubblicazione: (2024)
di: Feng, Jie, et al.
Pubblicazione: (2024)
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
Towards Robust Evaluation: A Comprehensive Taxonomy of Datasets and Metrics for Open Domain Question Answering in the Era of Large Language Models
di: Srivastava, Akchay, et al.
Pubblicazione: (2024)
di: Srivastava, Akchay, et al.
Pubblicazione: (2024)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
di: Sharma, Manasi, et al.
Pubblicazione: (2025)
di: Sharma, Manasi, et al.
Pubblicazione: (2025)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
di: Baidya, Avinash, et al.
Pubblicazione: (2025)
di: Baidya, Avinash, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
di: Hashimoto, Wataru, et al.
Pubblicazione: (2024) -
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025) -
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
di: Park, Jin Hyun, et al.
Pubblicazione: (2025) -
IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
di: Mohanty, Shrestha, et al.
Pubblicazione: (2024) -
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)