Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chung, Isaac, Kerboua, Imene, Kardos, Marton, Solomatin, Roman, Enevoldsen, Kenneth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MIEB: Massive Image Embedding Benchmark
von: Xiao, Chenghao, et al.
Veröffentlicht: (2025)
von: Xiao, Chenghao, et al.
Veröffentlicht: (2025)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2024)
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2024)
YABLoCo: Yet Another Benchmark for Long Context Code Generation
von: Valeev, Aidar, et al.
Veröffentlicht: (2025)
von: Valeev, Aidar, et al.
Veröffentlicht: (2025)
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
von: Han, Hojae, et al.
Veröffentlicht: (2025)
von: Han, Hojae, et al.
Veröffentlicht: (2025)
Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks
von: Cao, Jialun, et al.
Veröffentlicht: (2025)
von: Cao, Jialun, et al.
Veröffentlicht: (2025)
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
von: Chen, Jialong, et al.
Veröffentlicht: (2026)
von: Chen, Jialong, et al.
Veröffentlicht: (2026)
Can Coding Agents Reproduce Findings in Computational Materials Science?
von: Huang, Ziyang, et al.
Veröffentlicht: (2026)
von: Huang, Ziyang, et al.
Veröffentlicht: (2026)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
von: Chen, Zaoyu, et al.
Veröffentlicht: (2025)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2025)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
von: Ciancone, Mathieu, et al.
Veröffentlicht: (2024)
von: Ciancone, Mathieu, et al.
Veröffentlicht: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
IndustryCode: A Benchmark for Industry Code Generation
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
von: Diddee, Harshita, et al.
Veröffentlicht: (2026)
von: Diddee, Harshita, et al.
Veröffentlicht: (2026)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
A Code Comprehension Benchmark for Large Language Models for Code
von: Havare, Jayant, et al.
Veröffentlicht: (2025)
von: Havare, Jayant, et al.
Veröffentlicht: (2025)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
von: Yang, Jialin, et al.
Veröffentlicht: (2025)
von: Yang, Jialin, et al.
Veröffentlicht: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation
von: Harsh, Reetu Raj, et al.
Veröffentlicht: (2026)
von: Harsh, Reetu Raj, et al.
Veröffentlicht: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
von: Yang, Jie, et al.
Veröffentlicht: (2026)
von: Yang, Jie, et al.
Veröffentlicht: (2026)
AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
von: Zan, Daoguang, et al.
Veröffentlicht: (2024)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
Towards an Understanding of Context Utilization in Code Intelligence
von: Wang, Yanlin, et al.
Veröffentlicht: (2025)
von: Wang, Yanlin, et al.
Veröffentlicht: (2025)
CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
von: Cheng, Junhang, et al.
Veröffentlicht: (2026)
von: Cheng, Junhang, et al.
Veröffentlicht: (2026)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
von: Chen, Simin, et al.
Veröffentlicht: (2025)
von: Chen, Simin, et al.
Veröffentlicht: (2025)
The Model Openness Framework: Promoting Completeness and Openness for Reproducibility, Transparency, and Usability in Artificial Intelligence
von: White, Matt, et al.
Veröffentlicht: (2024)
von: White, Matt, et al.
Veröffentlicht: (2024)
E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
von: Liu, Jingyao, et al.
Veröffentlicht: (2025)
von: Liu, Jingyao, et al.
Veröffentlicht: (2025)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
SWE-AGI: Benchmarking Specification-Driven Software Construction with MoonBit in the Era of Autonomous Agents
von: Zhang, Zhirui, et al.
Veröffentlicht: (2026)
von: Zhang, Zhirui, et al.
Veröffentlicht: (2026)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
von: Cai, Songcheng, et al.
Veröffentlicht: (2026)
von: Cai, Songcheng, et al.
Veröffentlicht: (2026)
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
von: Adamenko, Pavel, et al.
Veröffentlicht: (2025)
von: Adamenko, Pavel, et al.
Veröffentlicht: (2025)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings
von: Gurioli, Andrea, et al.
Veröffentlicht: (2025)
von: Gurioli, Andrea, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MIEB: Massive Image Embedding Benchmark
von: Xiao, Chenghao, et al.
Veröffentlicht: (2025) -
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2024) -
YABLoCo: Yet Another Benchmark for Long Context Code Generation
von: Valeev, Aidar, et al.
Veröffentlicht: (2025) -
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
von: Han, Hojae, et al.
Veröffentlicht: (2025) -
Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks
von: Cao, Jialun, et al.
Veröffentlicht: (2025)