LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
Fuente:
arXiv
Salvato in:
| Autore principale: | Schlangen, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
di: Bhavsar, Nidhir, et al.
Pubblicazione: (2024)
di: Bhavsar, Nidhir, et al.
Pubblicazione: (2024)
Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models
di: Kennington, Casey, et al.
Pubblicazione: (2025)
di: Kennington, Casey, et al.
Pubblicazione: (2025)
Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes
di: Li, Meng, et al.
Pubblicazione: (2025)
di: Li, Meng, et al.
Pubblicazione: (2025)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
di: Beyer, Anne, et al.
Pubblicazione: (2024)
di: Beyer, Anne, et al.
Pubblicazione: (2024)
A Dual-Axis Taxonomy of Knowledge Editing for LLMs: From Mechanisms to Functions
di: Salehoof, Amir Mohammad, et al.
Pubblicazione: (2025)
di: Salehoof, Amir Mohammad, et al.
Pubblicazione: (2025)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
di: Zhou, Chengliang, et al.
Pubblicazione: (2025)
di: Zhou, Chengliang, et al.
Pubblicazione: (2025)
Qworld: Question-Specific Evaluation Criteria for LLMs
di: Gao, Shanghua, et al.
Pubblicazione: (2026)
di: Gao, Shanghua, et al.
Pubblicazione: (2026)
A Geometric Taxonomy of Hallucinations in LLMs
di: Marín, Javier
Pubblicazione: (2026)
di: Marín, Javier
Pubblicazione: (2026)
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering
di: Arias-Duart, Anna, et al.
Pubblicazione: (2025)
di: Arias-Duart, Anna, et al.
Pubblicazione: (2025)
On the Credibility of Evaluating LLMs using Survey Questions
di: Libovický, Jindřich
Pubblicazione: (2026)
di: Libovický, Jindřich
Pubblicazione: (2026)
Automated Analysis of Learning Outcomes and Exam Questions Based on Bloom's Taxonomy
di: Kumar, Ramya, et al.
Pubblicazione: (2025)
di: Kumar, Ramya, et al.
Pubblicazione: (2025)
Unraveling SITT: Social Influence Technique Taxonomy and Detection with LLMs
di: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Pubblicazione: (2025)
di: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Pubblicazione: (2025)
Evaluating and Enhancing LLMs for Multi-turn Text-to-SQL with Multiple Question Types
di: Guo, Ziming, et al.
Pubblicazione: (2024)
di: Guo, Ziming, et al.
Pubblicazione: (2024)
What Generative Artificial Intelligence Means for Terminological Definitions
di: Martín, Antonio San
Pubblicazione: (2024)
di: Martín, Antonio San
Pubblicazione: (2024)
Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models
di: Wu, Suhang, et al.
Pubblicazione: (2025)
di: Wu, Suhang, et al.
Pubblicazione: (2025)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
di: Kim, Minsang, et al.
Pubblicazione: (2024)
di: Kim, Minsang, et al.
Pubblicazione: (2024)
Can LLMs Ask Good Questions?
di: Zhang, Yueheng, et al.
Pubblicazione: (2025)
di: Zhang, Yueheng, et al.
Pubblicazione: (2025)
CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis
di: Zhang, Xinyu, et al.
Pubblicazione: (2025)
di: Zhang, Xinyu, et al.
Pubblicazione: (2025)
Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis
di: Wei, Yifan, et al.
Pubblicazione: (2026)
di: Wei, Yifan, et al.
Pubblicazione: (2026)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
di: Balamurali, Sai Shridhar, et al.
Pubblicazione: (2025)
di: Balamurali, Sai Shridhar, et al.
Pubblicazione: (2025)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
di: Wei, Jianhui, et al.
Pubblicazione: (2025)
di: Wei, Jianhui, et al.
Pubblicazione: (2025)
It Takes Two: A Dual Stage Approach for Terminology-Aware Translation
di: Jaswal, Akshat Singh
Pubblicazione: (2025)
di: Jaswal, Akshat Singh
Pubblicazione: (2025)
Assessing the Capability of LLMs in Solving POSCOMP Questions
di: Viegas, Cayo, et al.
Pubblicazione: (2025)
di: Viegas, Cayo, et al.
Pubblicazione: (2025)
Draft-based Approximate Inference for LLMs
di: Galim, Kevin, et al.
Pubblicazione: (2025)
di: Galim, Kevin, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
di: Raimondi, Bianca, et al.
Pubblicazione: (2026)
di: Raimondi, Bianca, et al.
Pubblicazione: (2026)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
di: Myrzakhan, Aidar, et al.
Pubblicazione: (2024)
di: Myrzakhan, Aidar, et al.
Pubblicazione: (2024)
MedCT: A Clinical Terminology Graph for Generative AI Applications in Healthcare
di: Chen, Ye, et al.
Pubblicazione: (2025)
di: Chen, Ye, et al.
Pubblicazione: (2025)
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
di: Hakimov, Sherzod, et al.
Pubblicazione: (2024)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2024)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
di: Hakimov, Sherzod, et al.
Pubblicazione: (2026)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2026)
How Effective is GPT-4 Turbo in Generating School-Level Questions from Textbooks Based on Bloom's Revised Taxonomy?
di: Maity, Subhankar, et al.
Pubblicazione: (2024)
di: Maity, Subhankar, et al.
Pubblicazione: (2024)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
di: Nejadgholi, Isar, et al.
Pubblicazione: (2025)
di: Nejadgholi, Isar, et al.
Pubblicazione: (2025)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
di: Mustapha, Ahmad, et al.
Pubblicazione: (2024)
di: Mustapha, Ahmad, et al.
Pubblicazione: (2024)
Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives
di: Dhar, Ruchira, et al.
Pubblicazione: (2026)
di: Dhar, Ruchira, et al.
Pubblicazione: (2026)
Efficient Technical Term Translation: A Knowledge Distillation Approach for Parenthetical Terminology Translation
di: Myung, Jiyoon, et al.
Pubblicazione: (2024)
di: Myung, Jiyoon, et al.
Pubblicazione: (2024)
CompoST: A Benchmark for Analyzing the Ability of LLMs To Compositionally Interpret Questions in a QALD Setting
di: Schmidt, David Maria, et al.
Pubblicazione: (2025)
di: Schmidt, David Maria, et al.
Pubblicazione: (2025)
Fine-Tuning LLMs for Reliable Medical Question-Answering Services
di: Anaissi, Ali, et al.
Pubblicazione: (2024)
di: Anaissi, Ali, et al.
Pubblicazione: (2024)
An Automatic Question Usability Evaluation Toolkit
di: Moore, Steven, et al.
Pubblicazione: (2024)
di: Moore, Steven, et al.
Pubblicazione: (2024)
DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
di: Panda, Srikant, et al.
Pubblicazione: (2025)
di: Panda, Srikant, et al.
Pubblicazione: (2025)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
di: Hong, Zijin, et al.
Pubblicazione: (2025)
di: Hong, Zijin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
di: Bhavsar, Nidhir, et al.
Pubblicazione: (2024) -
Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models
di: Kennington, Casey, et al.
Pubblicazione: (2025) -
Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes
di: Li, Meng, et al.
Pubblicazione: (2025) -
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
di: Beyer, Anne, et al.
Pubblicazione: (2024) -
A Dual-Axis Taxonomy of Knowledge Editing for LLMs: From Mechanisms to Functions
di: Salehoof, Amir Mohammad, et al.
Pubblicazione: (2025)