HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Assadi, Adnan El, Chung, Isaac, Solomatin, Roman, Muennighoff, Niklas, Enevoldsen, Kenneth |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
di: Chung, Isaac, et al.
Pubblicazione: (2025)
di: Chung, Isaac, et al.
Pubblicazione: (2025)
MIEB: Massive Image Embedding Benchmark
di: Xiao, Chenghao, et al.
Pubblicazione: (2025)
di: Xiao, Chenghao, et al.
Pubblicazione: (2025)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2024)
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2024)
MAEB: Massive Audio Embedding Benchmark
di: Assadi, Adnan El, et al.
Pubblicazione: (2026)
di: Assadi, Adnan El, et al.
Pubblicazione: (2026)
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
di: Kostiuk, Yevhen, et al.
Pubblicazione: (2026)
di: Kostiuk, Yevhen, et al.
Pubblicazione: (2026)
AutoIntent: AutoML for Text Classification
di: Alekseev, Ilya, et al.
Pubblicazione: (2025)
di: Alekseev, Ilya, et al.
Pubblicazione: (2025)
Grounding Text Embeddings in Stakeholder Associations
di: Rystrøm, Jonathan, et al.
Pubblicazione: (2026)
di: Rystrøm, Jonathan, et al.
Pubblicazione: (2026)
Exposing Assumptions in AI Benchmarks through Cognitive Modelling
di: Rystrøm, Jonathan H., et al.
Pubblicazione: (2024)
di: Rystrøm, Jonathan H., et al.
Pubblicazione: (2024)
topicwizard -- a Modern, Model-agnostic Framework for Topic Model Visualization and Interpretation
di: Kardos, Márton, et al.
Pubblicazione: (2025)
di: Kardos, Márton, et al.
Pubblicazione: (2025)
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
di: Nielsen, Dan Saattrup, et al.
Pubblicazione: (2024)
di: Nielsen, Dan Saattrup, et al.
Pubblicazione: (2024)
DANSK and DaCy 2.6.0: Domain Generalization of Danish Named Entity Recognition
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2024)
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2024)
C-Pack: Packed Resources For General Chinese Embeddings
di: Xiao, Shitao, et al.
Pubblicazione: (2023)
di: Xiao, Shitao, et al.
Pubblicazione: (2023)
Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models
di: Deng, Ningyuan, et al.
Pubblicazione: (2025)
di: Deng, Ningyuan, et al.
Pubblicazione: (2025)
Continuous sentiment scores for literary and multilingual contexts
di: Lyngbaek, Laurits, et al.
Pubblicazione: (2025)
di: Lyngbaek, Laurits, et al.
Pubblicazione: (2025)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
ATEB: Evaluating and Improving Advanced NLP Tasks for Text Embedding Models
di: Han, Simeng, et al.
Pubblicazione: (2025)
di: Han, Simeng, et al.
Pubblicazione: (2025)
Is Sentiment Banana-Shaped? Exploring the Geometry and Portability of Sentiment Concept Vectors
di: Lyngbaek, Laurits, et al.
Pubblicazione: (2026)
di: Lyngbaek, Laurits, et al.
Pubblicazione: (2026)
Naturalistic measure of social norms alignment
di: Kostiuk, Yevhen, et al.
Pubblicazione: (2026)
di: Kostiuk, Yevhen, et al.
Pubblicazione: (2026)
MMTEB: Massive Multilingual Text Embedding Benchmark
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2025)
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2025)
Which one Performs Better? Wav2Vec or Whisper? Applying both in Badini Kurdish Speech to Text (BKSTT)
di: Adnan, Renas, et al.
Pubblicazione: (2025)
di: Adnan, Renas, et al.
Pubblicazione: (2025)
Improving General Text Embedding Model: Tackling Task Conflict and Data Imbalance through Model Merging
di: Li, Mingxin, et al.
Pubblicazione: (2024)
di: Li, Mingxin, et al.
Pubblicazione: (2024)
ChemTEB: Chemical Text Embedding Benchmark, an Overview of Embedding Models Performance & Efficiency on a Specific Domain
di: Kasmaee, Ali Shiraee, et al.
Pubblicazione: (2024)
di: Kasmaee, Ali Shiraee, et al.
Pubblicazione: (2024)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
di: Kim, Eunsu, et al.
Pubblicazione: (2024)
di: Kim, Eunsu, et al.
Pubblicazione: (2024)
Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
di: Babakhin, Yauhen, et al.
Pubblicazione: (2025)
di: Babakhin, Yauhen, et al.
Pubblicazione: (2025)
Simulating User Diversity in Task-Oriented Dialogue Systems using Large Language Models
di: Ahmad, Adnan, et al.
Pubblicazione: (2025)
di: Ahmad, Adnan, et al.
Pubblicazione: (2025)
SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Detection
di: Muhammad, Shamsuddeen Hassan, et al.
Pubblicazione: (2025)
di: Muhammad, Shamsuddeen Hassan, et al.
Pubblicazione: (2025)
Enhancing Embedding Performance through Large Language Model-based Text Enrichment and Rewriting
di: Harris, Nicholas, et al.
Pubblicazione: (2024)
di: Harris, Nicholas, et al.
Pubblicazione: (2024)
Detecting Sexism in German Online Newspaper Comments with Open-Source Text Embeddings (Team GDA, GermEval2024 Shared Task 1: GerMS-Detect, Subtasks 1 and 2, Closed Track)
di: Bremm, Florian, et al.
Pubblicazione: (2024)
di: Bremm, Florian, et al.
Pubblicazione: (2024)
Mind the Gap: The Divergence Between Human and LLM-Generated Tasks
di: Lu, Yi-Long, et al.
Pubblicazione: (2025)
di: Lu, Yi-Long, et al.
Pubblicazione: (2025)
Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies
di: Tao, Chaofan, et al.
Pubblicazione: (2024)
di: Tao, Chaofan, et al.
Pubblicazione: (2024)
Towards Unified Task Embeddings Across Multiple Models: Bridging the Gap for Prompt-Based Large Language Models and Beyond
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
RegMix: Data Mixture as Regression for Language Model Pre-training
di: Liu, Qian, et al.
Pubblicazione: (2024)
di: Liu, Qian, et al.
Pubblicazione: (2024)
Linguistic and Embedding-Based Profiling of Texts generated by Humans and Large Language Models
di: Zanotto, Sergio E., et al.
Pubblicazione: (2025)
di: Zanotto, Sergio E., et al.
Pubblicazione: (2025)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
di: Longpre, Shayne, et al.
Pubblicazione: (2025)
di: Longpre, Shayne, et al.
Pubblicazione: (2025)
Bias in Text Embedding Models
di: Rakivnenko, Vasyl, et al.
Pubblicazione: (2024)
di: Rakivnenko, Vasyl, et al.
Pubblicazione: (2024)
Cross-Lingual Stability of LLM Judges Under Controlled Generation: Evidence from Finno-Ugric Languages
di: Chung, Isaac, et al.
Pubblicazione: (2026)
di: Chung, Isaac, et al.
Pubblicazione: (2026)
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
di: Zhang, Yanzhao, et al.
Pubblicazione: (2025)
di: Zhang, Yanzhao, et al.
Pubblicazione: (2025)
AIxcellent Vibes at GermEval 2025 Shared Task on Candy Speech Detection: Improving Model Performance by Span-Level Training
di: Thelen, Christian Rene, et al.
Pubblicazione: (2025)
di: Thelen, Christian Rene, et al.
Pubblicazione: (2025)
Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings
di: Icard, Benjamin, et al.
Pubblicazione: (2026)
di: Icard, Benjamin, et al.
Pubblicazione: (2026)
TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
di: Ezerceli, Özay, et al.
Pubblicazione: (2025)
di: Ezerceli, Özay, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
di: Chung, Isaac, et al.
Pubblicazione: (2025) -
MIEB: Massive Image Embedding Benchmark
di: Xiao, Chenghao, et al.
Pubblicazione: (2025) -
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2024) -
MAEB: Massive Audio Embedding Benchmark
di: Assadi, Adnan El, et al.
Pubblicazione: (2026) -
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
di: Kostiuk, Yevhen, et al.
Pubblicazione: (2026)