LM-PUB-QUIZ: A Comprehensive Framework for Zero-Shot Evaluation of Relational Knowledge in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ploner, Max, Wiland, Jacek, Pohl, Sebastian, Akbik, Alan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models
von: Wiland, Jacek, et al.
Veröffentlicht: (2024)
von: Wiland, Jacek, et al.
Veröffentlicht: (2024)
Towards a Principled Evaluation of Knowledge Editors
von: Pohl, Sebastian, et al.
Veröffentlicht: (2025)
von: Pohl, Sebastian, et al.
Veröffentlicht: (2025)
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
von: Christoph, Daniel, et al.
Veröffentlicht: (2025)
von: Christoph, Daniel, et al.
Veröffentlicht: (2025)
TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification Tasks
von: Garbas, Lukas, et al.
Veröffentlicht: (2024)
von: Garbas, Lukas, et al.
Veröffentlicht: (2024)
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
von: Golde, Jonas, et al.
Veröffentlicht: (2024)
von: Golde, Jonas, et al.
Veröffentlicht: (2024)
Self-Aware Knowledge Probing: Evaluating Language Models' Relational Knowledge through Confidence Calibration
von: Kissling, Christopher, et al.
Veröffentlicht: (2026)
von: Kissling, Christopher, et al.
Veröffentlicht: (2026)
Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs
von: Williams, Tristan, et al.
Veröffentlicht: (2026)
von: Williams, Tristan, et al.
Veröffentlicht: (2026)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation
von: Rücker, Susanna, et al.
Veröffentlicht: (2025)
von: Rücker, Susanna, et al.
Veröffentlicht: (2025)
SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2024)
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2024)
Large-Scale Label Interpretation Learning for Few-Shot Named Entity Recognition
von: Golde, Jonas, et al.
Veröffentlicht: (2024)
von: Golde, Jonas, et al.
Veröffentlicht: (2024)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
von: Haller, Patrick, et al.
Veröffentlicht: (2024)
von: Haller, Patrick, et al.
Veröffentlicht: (2024)
What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2026)
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2026)
Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
von: Toporkov, Olia, et al.
Veröffentlicht: (2025)
von: Toporkov, Olia, et al.
Veröffentlicht: (2025)
LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
von: Han, Chi, et al.
Veröffentlicht: (2023)
von: Han, Chi, et al.
Veröffentlicht: (2023)
Fundus: A Simple-to-Use News Scraper Optimized for High Quality Extractions
von: Dallabetta, Max, et al.
Veröffentlicht: (2024)
von: Dallabetta, Max, et al.
Veröffentlicht: (2024)
DecompressionLM: Deterministic, Diagnostic, and Zero-Shot Concept Graph Extraction from Language Models
von: Hong, Zhaochen, et al.
Veröffentlicht: (2026)
von: Hong, Zhaochen, et al.
Veröffentlicht: (2026)
What Matters When Building Universal Multilingual Named Entity Recognition Models?
von: Golde, Jonas, et al.
Veröffentlicht: (2026)
von: Golde, Jonas, et al.
Veröffentlicht: (2026)
PUB: Plot Understanding Benchmark and Dataset for Evaluating Large Language Models on Synthetic Visual Data Interpretation
von: Pawelec, Aneta, et al.
Veröffentlicht: (2024)
von: Pawelec, Aneta, et al.
Veröffentlicht: (2024)
A Comprehensive Evaluation of Semantic Relation Knowledge of Pretrained Language Models and Humans
von: Cao, Zhihan, et al.
Veröffentlicht: (2024)
von: Cao, Zhihan, et al.
Veröffentlicht: (2024)
OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
von: Lee, Unggi, et al.
Veröffentlicht: (2026)
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
von: Golde, Jonas, et al.
Veröffentlicht: (2025)
von: Golde, Jonas, et al.
Veröffentlicht: (2025)
ZeroLM: Data-Free Transformer Architecture Search for Language Models
von: Chen, Zhen-Song, et al.
Veröffentlicht: (2025)
von: Chen, Zhen-Song, et al.
Veröffentlicht: (2025)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
von: Muhamed, Aashiq
Veröffentlicht: (2025)
von: Muhamed, Aashiq
Veröffentlicht: (2025)
zrLLM: Zero-Shot Relational Learning on Temporal Knowledge Graphs with Large Language Models
von: Ding, Zifeng, et al.
Veröffentlicht: (2023)
von: Ding, Zifeng, et al.
Veröffentlicht: (2023)
MastermindEval: A Simple But Scalable Reasoning Benchmark
von: Golde, Jonas, et al.
Veröffentlicht: (2025)
von: Golde, Jonas, et al.
Veröffentlicht: (2025)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
von: Schulte, David, et al.
Veröffentlicht: (2024)
von: Schulte, David, et al.
Veröffentlicht: (2024)
Grasping the Essentials: Tailoring Large Language Models for Zero-Shot Relation Extraction
von: Zhou, Sizhe, et al.
Veröffentlicht: (2024)
von: Zhou, Sizhe, et al.
Veröffentlicht: (2024)
MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora
von: Feng, Tao, et al.
Veröffentlicht: (2026)
von: Feng, Tao, et al.
Veröffentlicht: (2026)
Question Decomposition for Retrieval-Augmented Generation
von: Ammann, Paul J. L., et al.
Veröffentlicht: (2025)
von: Ammann, Paul J. L., et al.
Veröffentlicht: (2025)
What the Weight?! A Unified Framework for Zero-Shot Knowledge Composition
von: Holtermann, Carolin, et al.
Veröffentlicht: (2024)
von: Holtermann, Carolin, et al.
Veröffentlicht: (2024)
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
von: Merdjanovska, Elena, et al.
Veröffentlicht: (2024)
von: Merdjanovska, Elena, et al.
Veröffentlicht: (2024)
Decoding Knowledge in Large Language Models: A Framework for Categorization and Comprehension
von: Fang, Yanbo, et al.
Veröffentlicht: (2025)
von: Fang, Yanbo, et al.
Veröffentlicht: (2025)
A Study on Building Efficient Zero-Shot Relation Extraction Models
von: Thomas, Hugo, et al.
Veröffentlicht: (2026)
von: Thomas, Hugo, et al.
Veröffentlicht: (2026)
Medical Coding with Biomedical Transformer Ensembles and Zero/Few-shot Learning
von: Ziletti, Angelo, et al.
Veröffentlicht: (2022)
von: Ziletti, Angelo, et al.
Veröffentlicht: (2022)
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities
von: Sravanthi, Settaluri Lakshmi, et al.
Veröffentlicht: (2024)
von: Sravanthi, Settaluri Lakshmi, et al.
Veröffentlicht: (2024)
KEDRec-LM: A Knowledge-distilled Explainable Drug Recommendation Large Language Model
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
von: Lin, Yujie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models
von: Wiland, Jacek, et al.
Veröffentlicht: (2024) -
Towards a Principled Evaluation of Knowledge Editors
von: Pohl, Sebastian, et al.
Veröffentlicht: (2025) -
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
von: Christoph, Daniel, et al.
Veröffentlicht: (2025) -
TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification Tasks
von: Garbas, Lukas, et al.
Veröffentlicht: (2024) -
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
von: Golde, Jonas, et al.
Veröffentlicht: (2024)