German Text Embedding Clustering Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Wehrli, Silvan, Arnrich, Bert, Irrgang, Christopher |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How to Measure the Intelligence of Large Language Models?
por: Körber, Nils, et al.
Publicado: (2024)
por: Körber, Nils, et al.
Publicado: (2024)
MMTEB: Massive Multilingual Text Embedding Benchmark
por: Enevoldsen, Kenneth, et al.
Publicado: (2025)
por: Enevoldsen, Kenneth, et al.
Publicado: (2025)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
por: Enevoldsen, Kenneth, et al.
Publicado: (2024)
por: Enevoldsen, Kenneth, et al.
Publicado: (2024)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
por: Pham, Loc, et al.
Publicado: (2025)
por: Pham, Loc, et al.
Publicado: (2025)
TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
por: Cao, Yang, et al.
Publicado: (2025)
por: Cao, Yang, et al.
Publicado: (2025)
Text-ADBench: Text Anomaly Detection Benchmark based on LLMs Embedding
por: Xiao, Feng, et al.
Publicado: (2025)
por: Xiao, Feng, et al.
Publicado: (2025)
ClustEm4Ano: Clustering Text Embeddings of Nominal Textual Attributes for Microdata Anonymization
por: Aufschläger, Robert, et al.
Publicado: (2024)
por: Aufschläger, Robert, et al.
Publicado: (2024)
Bias in Text Embedding Models
por: Rakivnenko, Vasyl, et al.
Publicado: (2024)
por: Rakivnenko, Vasyl, et al.
Publicado: (2024)
Classification of Human- and AI-Generated Texts for English, French, German, and Spanish
por: Schaaff, Kristina, et al.
Publicado: (2023)
por: Schaaff, Kristina, et al.
Publicado: (2023)
LLMs Enable Bag-of-Texts Representations for Short-Text Clustering
por: Lin, I-Fan, et al.
Publicado: (2025)
por: Lin, I-Fan, et al.
Publicado: (2025)
Contrastive Learning Subspace for Text Clustering
por: Yong, Qian, et al.
Publicado: (2024)
por: Yong, Qian, et al.
Publicado: (2024)
Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
por: Li, Shiyu, et al.
Publicado: (2025)
por: Li, Shiyu, et al.
Publicado: (2025)
EmbeddingGemma: Powerful and Lightweight Text Representations
por: Vera, Henrique Schechter, et al.
Publicado: (2025)
por: Vera, Henrique Schechter, et al.
Publicado: (2025)
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
por: Abdoli, Sajjad, et al.
Publicado: (2026)
por: Abdoli, Sajjad, et al.
Publicado: (2026)
Estimating Text Similarity based on Semantic Concept Embeddings
por: der Brück, Tim vor, et al.
Publicado: (2024)
por: der Brück, Tim vor, et al.
Publicado: (2024)
Interleaving Text and Number Embeddings to Solve Mathemathics Problems
por: Alberts, Marvin, et al.
Publicado: (2024)
por: Alberts, Marvin, et al.
Publicado: (2024)
Linearly-Interpretable Concept Embedding Models for Text Analysis
por: De Santis, Francesco, et al.
Publicado: (2024)
por: De Santis, Francesco, et al.
Publicado: (2024)
Grounding Text Embeddings in Stakeholder Associations
por: Rystrøm, Jonathan, et al.
Publicado: (2026)
por: Rystrøm, Jonathan, et al.
Publicado: (2026)
Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
por: Tsai, Yu-Che, et al.
Publicado: (2025)
por: Tsai, Yu-Che, et al.
Publicado: (2025)
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
por: Nagl, Sebastian, et al.
Publicado: (2026)
por: Nagl, Sebastian, et al.
Publicado: (2026)
SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applications
por: Lansiaux, Edouard, et al.
Publicado: (2025)
por: Lansiaux, Edouard, et al.
Publicado: (2025)
Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models
por: Deng, Ningyuan, et al.
Publicado: (2025)
por: Deng, Ningyuan, et al.
Publicado: (2025)
Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models
por: Chen, Kexin, et al.
Publicado: (2025)
por: Chen, Kexin, et al.
Publicado: (2025)
VISLA Benchmark: Evaluating Embedding Sensitivity to Semantic and Lexical Alterations
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities
por: Gao, Yingqiang, et al.
Publicado: (2025)
por: Gao, Yingqiang, et al.
Publicado: (2025)
Semantic-Driven Topic Modeling Using Transformer-Based Embeddings and Clustering Algorithms
por: Mersha, Melkamu Abay, et al.
Publicado: (2024)
por: Mersha, Melkamu Abay, et al.
Publicado: (2024)
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
por: Macko, Dominik, et al.
Publicado: (2024)
por: Macko, Dominik, et al.
Publicado: (2024)
Gecko: Versatile Text Embeddings Distilled from Large Language Models
por: Lee, Jinhyuk, et al.
Publicado: (2024)
por: Lee, Jinhyuk, et al.
Publicado: (2024)
Nomic Embed: Training a Reproducible Long Context Text Embedder
por: Nussbaum, Zach, et al.
Publicado: (2024)
por: Nussbaum, Zach, et al.
Publicado: (2024)
REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning
por: Lee, Seungmin, et al.
Publicado: (2026)
por: Lee, Seungmin, et al.
Publicado: (2026)
Text-as-Signal: Quantitative Semantic Scoring with Embeddings, Logprobs, and Noise Reduction
por: Moreira, Hugo
Publicado: (2026)
por: Moreira, Hugo
Publicado: (2026)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
por: Nguyen, Van Bach, et al.
Publicado: (2024)
por: Nguyen, Van Bach, et al.
Publicado: (2024)
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
por: Merrick, Luke, et al.
Publicado: (2024)
por: Merrick, Luke, et al.
Publicado: (2024)
Cequel: Cost-Effective Querying of Large Language Models for Text Clustering
por: Wang, Hongtao, et al.
Publicado: (2025)
por: Wang, Hongtao, et al.
Publicado: (2025)
Analysis of Utterance Embeddings and Clustering Methods Related to Intent Induction for Task-Oriented Dialogue
por: Park, Jeiyoon, et al.
Publicado: (2022)
por: Park, Jeiyoon, et al.
Publicado: (2022)
STAGE: Simplified Text-Attributed Graph Embeddings Using Pre-trained LLMs
por: Zolnai-Lucas, Aaron, et al.
Publicado: (2024)
por: Zolnai-Lucas, Aaron, et al.
Publicado: (2024)
Piccolo2: General Text Embedding with Multi-task Hybrid Loss Training
por: Huang, Junqin, et al.
Publicado: (2024)
por: Huang, Junqin, et al.
Publicado: (2024)
TEST: Text Prototype Aligned Embedding to Activate LLM's Ability for Time Series
por: Sun, Chenxi, et al.
Publicado: (2023)
por: Sun, Chenxi, et al.
Publicado: (2023)
QIME: Constructing Interpretable Medical Text Embeddings via Ontology-Grounded Questions
por: Tang, Yixuan, et al.
Publicado: (2026)
por: Tang, Yixuan, et al.
Publicado: (2026)
Responsible Diffusion Models via Constraining Text Embeddings within Safe Regions
por: Li, Zhiwen, et al.
Publicado: (2025)
por: Li, Zhiwen, et al.
Publicado: (2025)
Ejemplares similares
-
How to Measure the Intelligence of Large Language Models?
por: Körber, Nils, et al.
Publicado: (2024) -
MMTEB: Massive Multilingual Text Embedding Benchmark
por: Enevoldsen, Kenneth, et al.
Publicado: (2025) -
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
por: Enevoldsen, Kenneth, et al.
Publicado: (2024) -
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
por: Pham, Loc, et al.
Publicado: (2025) -
TAD-Bench: A Comprehensive Benchmark for Embedding-Based Text Anomaly Detection
por: Cao, Yang, et al.
Publicado: (2025)