SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
Fuente:
arXiv
Saved in:
| Main Authors: | Aynetdinov, Ansar, Akbik, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pre-Training Curriculum for Multi-Token Prediction in Language Models
by: Aynetdinov, Ansar, et al.
Published: (2025)
by: Aynetdinov, Ansar, et al.
Published: (2025)
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
by: Aynetdinov, Ansar, et al.
Published: (2026)
by: Aynetdinov, Ansar, et al.
Published: (2026)
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
by: Merdjanovska, Elena, et al.
Published: (2024)
by: Merdjanovska, Elena, et al.
Published: (2024)
Don't Mesh with Me: Generating Constructive Solid Geometry Instead of Meshes by Fine-Tuning a Code-Generation LLM
by: Mews, Maximilian, et al.
Published: (2024)
by: Mews, Maximilian, et al.
Published: (2024)
Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation
by: Rücker, Susanna, et al.
Published: (2025)
by: Rücker, Susanna, et al.
Published: (2025)
DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity
by: Zheng, Kaijie, et al.
Published: (2026)
by: Zheng, Kaijie, et al.
Published: (2026)
TexIm FAST: Text-to-Image Representation for Semantic Similarity Evaluation using Transformers
by: Ansar, Wazib, et al.
Published: (2024)
by: Ansar, Wazib, et al.
Published: (2024)
Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs
by: Williams, Tristan, et al.
Published: (2026)
by: Williams, Tristan, et al.
Published: (2026)
NLU-STR at SemEval-2024 Task 1: Generative-based Augmentation and Encoder-based Scoring for Semantic Textual Relatedness
by: Malaysha, Sanad, et al.
Published: (2024)
by: Malaysha, Sanad, et al.
Published: (2024)
Towards a Principled Evaluation of Knowledge Editors
by: Pohl, Sebastian, et al.
Published: (2025)
by: Pohl, Sebastian, et al.
Published: (2025)
BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models
by: Wiland, Jacek, et al.
Published: (2024)
by: Wiland, Jacek, et al.
Published: (2024)
Self-Aware Knowledge Probing: Evaluating Language Models' Relational Knowledge through Confidence Calibration
by: Kissling, Christopher, et al.
Published: (2026)
by: Kissling, Christopher, et al.
Published: (2026)
Linguistically Conditioned Semantic Textual Similarity
by: Tu, Jingxuan, et al.
Published: (2024)
by: Tu, Jingxuan, et al.
Published: (2024)
KurdSTS: The Kurdish Semantic Textual Similarity
by: Abdullah, Abdulhady Abas, et al.
Published: (2025)
by: Abdullah, Abdulhady Abas, et al.
Published: (2025)
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages
by: Ousidhoum, Nedjma, et al.
Published: (2024)
by: Ousidhoum, Nedjma, et al.
Published: (2024)
SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages
by: Ousidhoum, Nedjma, et al.
Published: (2024)
by: Ousidhoum, Nedjma, et al.
Published: (2024)
Explainable Semantic Textual Similarity via Dissimilar Span Detection
by: Lozano, Diego Miguel, et al.
Published: (2026)
by: Lozano, Diego Miguel, et al.
Published: (2026)
LM-PUB-QUIZ: A Comprehensive Framework for Zero-Shot Evaluation of Relational Knowledge in Language Models
by: Ploner, Max, et al.
Published: (2024)
by: Ploner, Max, et al.
Published: (2024)
Sharif-STR at SemEval-2024 Task 1: Transformer as a Regression Model for Fine-Grained Scoring of Textual Semantic Relations
by: Ebrahimi, Seyedeh Fatemeh, et al.
Published: (2024)
by: Ebrahimi, Seyedeh Fatemeh, et al.
Published: (2024)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
by: Haller, Patrick, et al.
Published: (2024)
by: Haller, Patrick, et al.
Published: (2024)
Large-Scale Label Interpretation Learning for Few-Shot Named Entity Recognition
by: Golde, Jonas, et al.
Published: (2024)
by: Golde, Jonas, et al.
Published: (2024)
TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification Tasks
by: Garbas, Lukas, et al.
Published: (2024)
by: Garbas, Lukas, et al.
Published: (2024)
What Matters When Building Universal Multilingual Named Entity Recognition Models?
by: Golde, Jonas, et al.
Published: (2026)
by: Golde, Jonas, et al.
Published: (2026)
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
by: Golde, Jonas, et al.
Published: (2025)
by: Golde, Jonas, et al.
Published: (2025)
Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
by: Toporkov, Olia, et al.
Published: (2025)
by: Toporkov, Olia, et al.
Published: (2025)
MasonTigers at SemEval-2024 Task 1: An Ensemble Approach for Semantic Textual Relatedness
by: Goswami, Dhiman, et al.
Published: (2024)
by: Goswami, Dhiman, et al.
Published: (2024)
Are ELECTRA's Sentence Embeddings Beyond Repair? The Case of Semantic Textual Similarity
by: Rep, Ivan, et al.
Published: (2024)
by: Rep, Ivan, et al.
Published: (2024)
Pcc-tuning: Breaking the Contrastive Learning Ceiling in Semantic Textual Similarity
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
by: Christoph, Daniel, et al.
Published: (2025)
by: Christoph, Daniel, et al.
Published: (2025)
AAdaM at SemEval-2024 Task 1: Augmentation and Adaptation for Multilingual Semantic Textual Relatedness
by: Zhang, Miaoran, et al.
Published: (2024)
by: Zhang, Miaoran, et al.
Published: (2024)
Multilingual Evaluation of Semantic Textual Relatedness
by: Endait, Sharvi, et al.
Published: (2024)
by: Endait, Sharvi, et al.
Published: (2024)
Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
by: Golde, Jonas, et al.
Published: (2023)
by: Golde, Jonas, et al.
Published: (2023)
CASE -- Condition-Aware Sentence Embeddings for Conditional Semantic Textual Similarity Measurement
by: Zhang, Gaifan, et al.
Published: (2025)
by: Zhang, Gaifan, et al.
Published: (2025)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
by: Schulte, David, et al.
Published: (2024)
by: Schulte, David, et al.
Published: (2024)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
PESTS: Persian_English Cross Lingual Corpus for Semantic Textual Similarity
by: Abdous, Mohammad, et al.
Published: (2023)
by: Abdous, Mohammad, et al.
Published: (2023)
Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Cross-lingual Transfer or Machine Translation? On Data Augmentation for Monolingual Semantic Textual Similarity
by: Hoshino, Sho, et al.
Published: (2024)
by: Hoshino, Sho, et al.
Published: (2024)
Advances and Challenges in Semantic Textual Similarity: A Comprehensive Survey
by: Kumar, Lokendra, et al.
Published: (2025)
by: Kumar, Lokendra, et al.
Published: (2025)
Similar Items
-
Pre-Training Curriculum for Multi-Token Prediction in Language Models
by: Aynetdinov, Ansar, et al.
Published: (2025) -
Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
by: Aynetdinov, Ansar, et al.
Published: (2026) -
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
by: Merdjanovska, Elena, et al.
Published: (2024) -
Don't Mesh with Me: Generating Constructive Solid Geometry Instead of Meshes by Fine-Tuning a Code-Generation LLM
by: Mews, Maximilian, et al.
Published: (2024) -
Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation
by: Rücker, Susanna, et al.
Published: (2025)