The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
Fuente:
arXiv
Saved in:
| Main Authors: | Enevoldsen, Kenneth, Kardos, Márton, Muennighoff, Niklas, Nielbo, Kristoffer Laigaard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
topicwizard -- a Modern, Model-agnostic Framework for Topic Model Visualization and Interpretation
by: Kardos, Márton, et al.
Published: (2025)
by: Kardos, Márton, et al.
Published: (2025)
Improving reasoning at inference time via uncertainty minimisation
by: Legrand, Nicolas, et al.
Published: (2026)
by: Legrand, Nicolas, et al.
Published: (2026)
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
by: Chung, Isaac, et al.
Published: (2025)
by: Chung, Isaac, et al.
Published: (2025)
Naturalistic measure of social norms alignment
by: Kostiuk, Yevhen, et al.
Published: (2026)
by: Kostiuk, Yevhen, et al.
Published: (2026)
MIEB: Massive Image Embedding Benchmark
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
$S^3$ -- Semantic Signal Separation
by: Kardos, Márton, et al.
Published: (2024)
by: Kardos, Márton, et al.
Published: (2024)
MMTEB: Massive Multilingual Text Embedding Benchmark
by: Enevoldsen, Kenneth, et al.
Published: (2025)
by: Enevoldsen, Kenneth, et al.
Published: (2025)
HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
by: Assadi, Adnan El, et al.
Published: (2025)
by: Assadi, Adnan El, et al.
Published: (2025)
Grounding Text Embeddings in Stakeholder Associations
by: Rystrøm, Jonathan, et al.
Published: (2026)
by: Rystrøm, Jonathan, et al.
Published: (2026)
MAEB: Massive Audio Embedding Benchmark
by: Assadi, Adnan El, et al.
Published: (2026)
by: Assadi, Adnan El, et al.
Published: (2026)
Topeax -- An Improved Clustering Topic Model with Density Peak Detection and Lexical-Semantic Term Importance
by: Kardos, Márton
Published: (2026)
by: Kardos, Márton
Published: (2026)
Dynaword: From One-shot to Continuously Developed Datasets
by: Enevoldsen, Kenneth, et al.
Published: (2025)
by: Enevoldsen, Kenneth, et al.
Published: (2025)
Exposing Assumptions in AI Benchmarks through Cognitive Modelling
by: Rystrøm, Jonathan H., et al.
Published: (2024)
by: Rystrøm, Jonathan H., et al.
Published: (2024)
Monolingual or Multilingual Instruction Tuning: Which Makes a Better Alpaca
by: Chen, Pinzhen, et al.
Published: (2023)
by: Chen, Pinzhen, et al.
Published: (2023)
Continuous sentiment scores for literary and multilingual contexts
by: Lyngbaek, Laurits, et al.
Published: (2025)
by: Lyngbaek, Laurits, et al.
Published: (2025)
C-Pack: Packed Resources For General Chinese Embeddings
by: Xiao, Shitao, et al.
Published: (2023)
by: Xiao, Shitao, et al.
Published: (2023)
TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
German Text Embedding Clustering Benchmark
by: Wehrli, Silvan, et al.
Published: (2024)
by: Wehrli, Silvan, et al.
Published: (2024)
Team QUST at SemEval-2024 Task 8: A Comprehensive Study of Monolingual and Multilingual Approaches for Detecting AI-generated Text
by: Xu, Xiaoman, et al.
Published: (2024)
by: Xu, Xiaoman, et al.
Published: (2024)
Are Chatbots Reliable Text Annotators? Sometimes
by: Kristensen-McLachlan, Ross Deans, et al.
Published: (2023)
by: Kristensen-McLachlan, Ross Deans, et al.
Published: (2023)
Is Sentiment Banana-Shaped? Exploring the Geometry and Portability of Sentiment Concept Vectors
by: Lyngbaek, Laurits, et al.
Published: (2026)
by: Lyngbaek, Laurits, et al.
Published: (2026)
Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark
by: Mastrokostas, Charalampos, et al.
Published: (2026)
by: Mastrokostas, Charalampos, et al.
Published: (2026)
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
by: Zhang, Ziyin, et al.
Published: (2026)
by: Zhang, Ziyin, et al.
Published: (2026)
Text Embedding Inversion Security for Multilingual Language Models
by: Chen, Yiyi, et al.
Published: (2024)
by: Chen, Yiyi, et al.
Published: (2024)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
by: Pham, Loc, et al.
Published: (2025)
by: Pham, Loc, et al.
Published: (2025)
Multilingual Information Retrieval with a Monolingual Knowledge Base
by: Zhuang, Yingying, et al.
Published: (2025)
by: Zhuang, Yingying, et al.
Published: (2025)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings
by: Ueareeworakul, Pakorn, et al.
Published: (2025)
by: Ueareeworakul, Pakorn, et al.
Published: (2025)
Enhancing Multilingual Embeddings via Multi-Way Parallel Text Alignment
by: Fazili, Barah, et al.
Published: (2026)
by: Fazili, Barah, et al.
Published: (2026)
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
by: Nielsen, Dan Saattrup, et al.
Published: (2024)
by: Nielsen, Dan Saattrup, et al.
Published: (2024)
Bias in Text Embedding Models
by: Rakivnenko, Vasyl, et al.
Published: (2024)
by: Rakivnenko, Vasyl, et al.
Published: (2024)
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
by: Li, Shiyu, et al.
Published: (2025)
by: Li, Shiyu, et al.
Published: (2025)
Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
by: Tsai, Yu-Che, et al.
Published: (2025)
by: Tsai, Yu-Che, et al.
Published: (2025)
EmbeddingGemma: Powerful and Lightweight Text Representations
by: Vera, Henrique Schechter, et al.
Published: (2025)
by: Vera, Henrique Schechter, et al.
Published: (2025)
When Text Embedding Meets Large Language Model: A Comprehensive Survey
by: Nie, Zhijie, et al.
Published: (2024)
by: Nie, Zhijie, et al.
Published: (2024)
A Comprehensive Analysis of Static Word Embeddings for Turkish
by: Sarıtaş, Karahan, et al.
Published: (2024)
by: Sarıtaş, Karahan, et al.
Published: (2024)
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
by: Merrick, Luke, et al.
Published: (2024)
by: Merrick, Luke, et al.
Published: (2024)
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
by: Kostiuk, Yevhen, et al.
Published: (2026)
by: Kostiuk, Yevhen, et al.
Published: (2026)
Using Optimal Transport as Alignment Objective for fine-tuning Multilingual Contextualized Embeddings
by: Alqahtani, Sawsan, et al.
Published: (2021)
by: Alqahtani, Sawsan, et al.
Published: (2021)
Similar Items
-
topicwizard -- a Modern, Model-agnostic Framework for Topic Model Visualization and Interpretation
by: Kardos, Márton, et al.
Published: (2025) -
Improving reasoning at inference time via uncertainty minimisation
by: Legrand, Nicolas, et al.
Published: (2026) -
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
by: Chung, Isaac, et al.
Published: (2025) -
Naturalistic measure of social norms alignment
by: Kostiuk, Yevhen, et al.
Published: (2026) -
MIEB: Massive Image Embedding Benchmark
by: Xiao, Chenghao, et al.
Published: (2025)