German Text Embedding Clustering Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Wehrli, Silvan, Arnrich, Bert, Irrgang, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Measure the Intelligence of Large Language Models?
by: Körber, Nils, et al.
Published: (2024)
by: Körber, Nils, et al.
Published: (2024)
MMTEB: Massive Multilingual Text Embedding Benchmark
by: Enevoldsen, Kenneth, et al.
Published: (2025)
by: Enevoldsen, Kenneth, et al.
Published: (2025)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
by: Enevoldsen, Kenneth, et al.
Published: (2024)
by: Enevoldsen, Kenneth, et al.
Published: (2024)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
by: Pham, Loc, et al.
Published: (2025)
by: Pham, Loc, et al.
Published: (2025)
TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
ClustEm4Ano: Clustering Text Embeddings of Nominal Textual Attributes for Microdata Anonymization
by: Aufschläger, Robert, et al.
Published: (2024)
by: Aufschläger, Robert, et al.
Published: (2024)
Bias in Text Embedding Models
by: Rakivnenko, Vasyl, et al.
Published: (2024)
by: Rakivnenko, Vasyl, et al.
Published: (2024)
Classification of Human- and AI-Generated Texts for English, French, German, and Spanish
by: Schaaff, Kristina, et al.
Published: (2023)
by: Schaaff, Kristina, et al.
Published: (2023)
LLMs Enable Bag-of-Texts Representations for Short-Text Clustering
by: Lin, I-Fan, et al.
Published: (2025)
by: Lin, I-Fan, et al.
Published: (2025)
Contrastive Learning Subspace for Text Clustering
by: Yong, Qian, et al.
Published: (2024)
by: Yong, Qian, et al.
Published: (2024)
Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
by: Li, Shiyu, et al.
Published: (2025)
by: Li, Shiyu, et al.
Published: (2025)
EmbeddingGemma: Powerful and Lightweight Text Representations
by: Vera, Henrique Schechter, et al.
Published: (2025)
by: Vera, Henrique Schechter, et al.
Published: (2025)
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
by: Abdoli, Sajjad, et al.
Published: (2026)
by: Abdoli, Sajjad, et al.
Published: (2026)
Estimating Text Similarity based on Semantic Concept Embeddings
by: der Brück, Tim vor, et al.
Published: (2024)
by: der Brück, Tim vor, et al.
Published: (2024)
Interleaving Text and Number Embeddings to Solve Mathemathics Problems
by: Alberts, Marvin, et al.
Published: (2024)
by: Alberts, Marvin, et al.
Published: (2024)
Linearly-Interpretable Concept Embedding Models for Text Analysis
by: De Santis, Francesco, et al.
Published: (2024)
by: De Santis, Francesco, et al.
Published: (2024)
Grounding Text Embeddings in Stakeholder Associations
by: Rystrøm, Jonathan, et al.
Published: (2026)
by: Rystrøm, Jonathan, et al.
Published: (2026)
Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
by: Tsai, Yu-Che, et al.
Published: (2025)
by: Tsai, Yu-Che, et al.
Published: (2025)
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
by: Nagl, Sebastian, et al.
Published: (2026)
by: Nagl, Sebastian, et al.
Published: (2026)
SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
by: Lansiaux, Edouard, et al.
Published: (2025)
by: Lansiaux, Edouard, et al.
Published: (2025)
Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models
by: Deng, Ningyuan, et al.
Published: (2025)
by: Deng, Ningyuan, et al.
Published: (2025)
Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models
by: Chen, Kexin, et al.
Published: (2025)
by: Chen, Kexin, et al.
Published: (2025)
VISLA Benchmark: Evaluating Embedding Sensitivity to Semantic and Lexical Alterations
by: Dumpala, Sri Harsha, et al.
Published: (2024)
by: Dumpala, Sri Harsha, et al.
Published: (2024)
Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities
by: Gao, Yingqiang, et al.
Published: (2025)
by: Gao, Yingqiang, et al.
Published: (2025)
Semantic-Driven Topic Modeling Using Transformer-Based Embeddings and Clustering Algorithms
by: Mersha, Melkamu Abay, et al.
Published: (2024)
by: Mersha, Melkamu Abay, et al.
Published: (2024)
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
by: Macko, Dominik, et al.
Published: (2024)
by: Macko, Dominik, et al.
Published: (2024)
Gecko: Versatile Text Embeddings Distilled from Large Language Models
by: Lee, Jinhyuk, et al.
Published: (2024)
by: Lee, Jinhyuk, et al.
Published: (2024)
Nomic Embed: Training a Reproducible Long Context Text Embedder
by: Nussbaum, Zach, et al.
Published: (2024)
by: Nussbaum, Zach, et al.
Published: (2024)
REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning
by: Lee, Seungmin, et al.
Published: (2026)
by: Lee, Seungmin, et al.
Published: (2026)
Text-as-Signal: Quantitative Semantic Scoring with Embeddings, Logprobs, and Noise Reduction
by: Moreira, Hugo
Published: (2026)
by: Moreira, Hugo
Published: (2026)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
by: Nguyen, Van Bach, et al.
Published: (2024)
by: Nguyen, Van Bach, et al.
Published: (2024)
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
by: Merrick, Luke, et al.
Published: (2024)
by: Merrick, Luke, et al.
Published: (2024)
Cequel: Cost-Effective Querying of Large Language Models for Text Clustering
by: Wang, Hongtao, et al.
Published: (2025)
by: Wang, Hongtao, et al.
Published: (2025)
Analysis of Utterance Embeddings and Clustering Methods Related to Intent Induction for Task-Oriented Dialogue
by: Park, Jeiyoon, et al.
Published: (2022)
by: Park, Jeiyoon, et al.
Published: (2022)
STAGE: Simplified Text-Attributed Graph Embeddings Using Pre-trained LLMs
by: Zolnai-Lucas, Aaron, et al.
Published: (2024)
by: Zolnai-Lucas, Aaron, et al.
Published: (2024)
Piccolo2: General Text Embedding with Multi-task Hybrid Loss Training
by: Huang, Junqin, et al.
Published: (2024)
by: Huang, Junqin, et al.
Published: (2024)
TEST: Text Prototype Aligned Embedding to Activate LLM's Ability for Time Series
by: Sun, Chenxi, et al.
Published: (2023)
by: Sun, Chenxi, et al.
Published: (2023)
QIME: Constructing Interpretable Medical Text Embeddings via Ontology-Grounded Questions
by: Tang, Yixuan, et al.
Published: (2026)
by: Tang, Yixuan, et al.
Published: (2026)
Responsible Diffusion Models via Constraining Text Embeddings within Safe Regions
by: Li, Zhiwen, et al.
Published: (2025)
by: Li, Zhiwen, et al.
Published: (2025)
Similar Items
-
How to Measure the Intelligence of Large Language Models?
by: Körber, Nils, et al.
Published: (2024) -
MMTEB: Massive Multilingual Text Embedding Benchmark
by: Enevoldsen, Kenneth, et al.
Published: (2025) -
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
by: Enevoldsen, Kenneth, et al.
Published: (2024) -
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
by: Pham, Loc, et al.
Published: (2025) -
TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
by: Cao, Yang, et al.
Published: (2025)