Guardado en:
| Autores principales: | Pohl, Sebastian, Ploner, Max, Akbik, Alan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2507.05937 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LM-PUB-QUIZ: A Comprehensive Framework for Zero-Shot Evaluation of Relational Knowledge in Language Models
por: Ploner, Max, et al.
Publicado: (2024)
por: Ploner, Max, et al.
Publicado: (2024)
BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models
por: Wiland, Jacek, et al.
Publicado: (2024)
por: Wiland, Jacek, et al.
Publicado: (2024)
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
por: Christoph, Daniel, et al.
Publicado: (2025)
por: Christoph, Daniel, et al.
Publicado: (2025)
TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification Tasks
por: Garbas, Lukas, et al.
Publicado: (2024)
por: Garbas, Lukas, et al.
Publicado: (2024)
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
por: Golde, Jonas, et al.
Publicado: (2024)
por: Golde, Jonas, et al.
Publicado: (2024)
Self-Aware Knowledge Probing: Evaluating Language Models' Relational Knowledge through Confidence Calibration
por: Kissling, Christopher, et al.
Publicado: (2026)
por: Kissling, Christopher, et al.
Publicado: (2026)
Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation
por: Rücker, Susanna, et al.
Publicado: (2025)
por: Rücker, Susanna, et al.
Publicado: (2025)
SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
por: Aynetdinov, Ansar, et al.
Publicado: (2024)
por: Aynetdinov, Ansar, et al.
Publicado: (2024)
Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs
por: Williams, Tristan, et al.
Publicado: (2026)
por: Williams, Tristan, et al.
Publicado: (2026)
Fundus: A Simple-to-Use News Scraper Optimized for High Quality Extractions
por: Dallabetta, Max, et al.
Publicado: (2024)
por: Dallabetta, Max, et al.
Publicado: (2024)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
por: Aynetdinov, Ansar, et al.
Publicado: (2025)
por: Aynetdinov, Ansar, et al.
Publicado: (2025)
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
por: Golde, Jonas, et al.
Publicado: (2025)
por: Golde, Jonas, et al.
Publicado: (2025)
Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
por: Toporkov, Olia, et al.
Publicado: (2025)
por: Toporkov, Olia, et al.
Publicado: (2025)
What Matters When Building Universal Multilingual Named Entity Recognition Models?
por: Golde, Jonas, et al.
Publicado: (2026)
por: Golde, Jonas, et al.
Publicado: (2026)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
por: Haller, Patrick, et al.
Publicado: (2024)
por: Haller, Patrick, et al.
Publicado: (2024)
Large-Scale Label Interpretation Learning for Few-Shot Named Entity Recognition
por: Golde, Jonas, et al.
Publicado: (2024)
por: Golde, Jonas, et al.
Publicado: (2024)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
por: Haller, Patrick, et al.
Publicado: (2025)
por: Haller, Patrick, et al.
Publicado: (2025)
What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
por: Haller, Patrick, et al.
Publicado: (2025)
por: Haller, Patrick, et al.
Publicado: (2025)
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
por: Aynetdinov, Ansar, et al.
Publicado: (2026)
por: Aynetdinov, Ansar, et al.
Publicado: (2026)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
por: Schulte, David, et al.
Publicado: (2024)
por: Schulte, David, et al.
Publicado: (2024)
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
por: Merdjanovska, Elena, et al.
Publicado: (2024)
por: Merdjanovska, Elena, et al.
Publicado: (2024)
MastermindEval: A Simple But Scalable Reasoning Benchmark
por: Golde, Jonas, et al.
Publicado: (2025)
por: Golde, Jonas, et al.
Publicado: (2025)
Question Decomposition for Retrieval-Augmented Generation
por: Ammann, Paul J. L., et al.
Publicado: (2025)
por: Ammann, Paul J. L., et al.
Publicado: (2025)
Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
por: Golde, Jonas, et al.
Publicado: (2023)
por: Golde, Jonas, et al.
Publicado: (2023)
Medical Coding with Biomedical Transformer Ensembles and Zero/Few-shot Learning
por: Ziletti, Angelo, et al.
Publicado: (2022)
por: Ziletti, Angelo, et al.
Publicado: (2022)
HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization tools
por: Sänger, Mario, et al.
Publicado: (2024)
por: Sänger, Mario, et al.
Publicado: (2024)
Mind Your Format: Towards Consistent Evaluation of In-Context Learning Improvements
por: Voronov, Anton, et al.
Publicado: (2024)
por: Voronov, Anton, et al.
Publicado: (2024)
Model-Aware Tokenizer Transfer
por: Haltiuk, Mykola, et al.
Publicado: (2025)
por: Haltiuk, Mykola, et al.
Publicado: (2025)
Targum -- A Multilingual New Testament Translation Corpus
por: Rapacz, Maciej, et al.
Publicado: (2026)
por: Rapacz, Maciej, et al.
Publicado: (2026)
WilKE: Wise-Layer Knowledge Editor for Lifelong Knowledge Editing
por: Hu, Chenhui, et al.
Publicado: (2024)
por: Hu, Chenhui, et al.
Publicado: (2024)
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
por: Borkakoty, Hsuvas, et al.
Publicado: (2026)
por: Borkakoty, Hsuvas, et al.
Publicado: (2026)
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall
por: Yuan, Jiaqing, et al.
Publicado: (2024)
por: Yuan, Jiaqing, et al.
Publicado: (2024)
Towards Efficient LLMs Annealing with Principled Sample Selection
por: Xu, Yuanjian, et al.
Publicado: (2026)
por: Xu, Yuanjian, et al.
Publicado: (2026)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
por: Panigrahi, Abhishek, et al.
Publicado: (2025)
por: Panigrahi, Abhishek, et al.
Publicado: (2025)
A Principled Framework for Evaluating on Typologically Diverse Languages
por: Ploeger, Esther, et al.
Publicado: (2024)
por: Ploeger, Esther, et al.
Publicado: (2024)
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
por: Ku, Max, et al.
Publicado: (2023)
por: Ku, Max, et al.
Publicado: (2023)
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
por: Oh, Hongseok, et al.
Publicado: (2025)
por: Oh, Hongseok, et al.
Publicado: (2025)
Reverse Probing: Evaluating Knowledge Transfer via Finetuned Task Embeddings for Coreference Resolution
por: Anikina, Tatiana, et al.
Publicado: (2025)
por: Anikina, Tatiana, et al.
Publicado: (2025)
Do Language Models Encode Knowledge of Linguistic Constraint Violations?
por: Hardy, et al.
Publicado: (2026)
por: Hardy, et al.
Publicado: (2026)
PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
por: Haller, Patrick, et al.
Publicado: (2025)
por: Haller, Patrick, et al.
Publicado: (2025)
Ejemplares similares
-
LM-PUB-QUIZ: A Comprehensive Framework for Zero-Shot Evaluation of Relational Knowledge in Language Models
por: Ploner, Max, et al.
Publicado: (2024) -
BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models
por: Wiland, Jacek, et al.
Publicado: (2024) -
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
por: Christoph, Daniel, et al.
Publicado: (2025) -
TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification Tasks
por: Garbas, Lukas, et al.
Publicado: (2024) -
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
por: Golde, Jonas, et al.
Publicado: (2024)