Maintaining MTEB: Towards Long Term Usability and Reproducibility of Embedding Benchmarks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chung, Isaac, Kerboua, Imene, Kardos, Marton, Solomatin, Roman, Enevoldsen, Kenneth |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MIEB: Massive Image Embedding Benchmark
par: Xiao, Chenghao, et autres
Publié: (2025)
par: Xiao, Chenghao, et autres
Publié: (2025)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
par: Enevoldsen, Kenneth, et autres
Publié: (2024)
par: Enevoldsen, Kenneth, et autres
Publié: (2024)
YABLoCo: Yet Another Benchmark for Long Context Code Generation
par: Valeev, Aidar, et autres
Publié: (2025)
par: Valeev, Aidar, et autres
Publié: (2025)
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
par: Han, Hojae, et autres
Publié: (2025)
par: Han, Hojae, et autres
Publié: (2025)
Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks
par: Cao, Jialun, et autres
Publié: (2025)
par: Cao, Jialun, et autres
Publié: (2025)
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
par: Chen, Jialong, et autres
Publié: (2026)
par: Chen, Jialong, et autres
Publié: (2026)
Can Coding Agents Reproduce Findings in Computational Materials Science?
par: Huang, Ziyang, et autres
Publié: (2026)
par: Huang, Ziyang, et autres
Publié: (2026)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
par: Chen, Zaoyu, et autres
Publié: (2025)
par: Chen, Zaoyu, et autres
Publié: (2025)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
par: Orlanski, Gabriel, et autres
Publié: (2026)
par: Orlanski, Gabriel, et autres
Publié: (2026)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
par: Sonwane, Atharv, et autres
Publié: (2026)
par: Sonwane, Atharv, et autres
Publié: (2026)
MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
par: Ciancone, Mathieu, et autres
Publié: (2024)
par: Ciancone, Mathieu, et autres
Publié: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
par: Tu, Xinming, et autres
Publié: (2026)
par: Tu, Xinming, et autres
Publié: (2026)
IndustryCode: A Benchmark for Industry Code Generation
par: Zeng, Puyu, et autres
Publié: (2026)
par: Zeng, Puyu, et autres
Publié: (2026)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
par: Diddee, Harshita, et autres
Publié: (2026)
par: Diddee, Harshita, et autres
Publié: (2026)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
par: Zan, Daoguang, et autres
Publié: (2025)
par: Zan, Daoguang, et autres
Publié: (2025)
A Code Comprehension Benchmark for Large Language Models for Code
par: Havare, Jayant, et autres
Publié: (2025)
par: Havare, Jayant, et autres
Publié: (2025)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
par: Yang, Jialin, et autres
Publié: (2025)
par: Yang, Jialin, et autres
Publié: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
par: Jiang, Hongchao, et autres
Publié: (2025)
par: Jiang, Hongchao, et autres
Publié: (2025)
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation
par: Harsh, Reetu Raj, et autres
Publié: (2026)
par: Harsh, Reetu Raj, et autres
Publié: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
par: Yang, Jie, et autres
Publié: (2026)
par: Yang, Jie, et autres
Publié: (2026)
AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
par: Zhao, Haoyu, et autres
Publié: (2026)
par: Zhao, Haoyu, et autres
Publié: (2026)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
par: Jiang, Nan, et autres
Publié: (2024)
par: Jiang, Nan, et autres
Publié: (2024)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
par: Zhao, Bingchen, et autres
Publié: (2026)
par: Zhao, Bingchen, et autres
Publié: (2026)
SWE-bench-java: A GitHub Issue Resolving Benchmark for Java
par: Zan, Daoguang, et autres
Publié: (2024)
par: Zan, Daoguang, et autres
Publié: (2024)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
par: Lam, Man Ho, et autres
Publié: (2026)
par: Lam, Man Ho, et autres
Publié: (2026)
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
par: Zheng, Jiasheng, et autres
Publié: (2024)
par: Zheng, Jiasheng, et autres
Publié: (2024)
Towards an Understanding of Context Utilization in Code Intelligence
par: Wang, Yanlin, et autres
Publié: (2025)
par: Wang, Yanlin, et autres
Publié: (2025)
CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
par: Cheng, Junhang, et autres
Publié: (2026)
par: Cheng, Junhang, et autres
Publié: (2026)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
par: Zhuo, Terry Yue, et autres
Publié: (2024)
par: Zhuo, Terry Yue, et autres
Publié: (2024)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
par: Chen, Simin, et autres
Publié: (2025)
par: Chen, Simin, et autres
Publié: (2025)
The Model Openness Framework: Promoting Completeness and Openness for Reproducibility, Transparency, and Usability in Artificial Intelligence
par: White, Matt, et autres
Publié: (2024)
par: White, Matt, et autres
Publié: (2024)
E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
par: Liu, Jingyao, et autres
Publié: (2025)
par: Liu, Jingyao, et autres
Publié: (2025)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
SWE-AGI: Benchmarking Specification-Driven Software Construction with MoonBit in the Era of Autonomous Agents
par: Zhang, Zhirui, et autres
Publié: (2026)
par: Zhang, Zhirui, et autres
Publié: (2026)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
par: Wang, Yanli, et autres
Publié: (2024)
par: Wang, Yanli, et autres
Publié: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
par: Cai, Songcheng, et autres
Publié: (2026)
par: Cai, Songcheng, et autres
Publié: (2026)
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
par: Adamenko, Pavel, et autres
Publié: (2025)
par: Adamenko, Pavel, et autres
Publié: (2025)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
par: Yan, Weixiang, et autres
Publié: (2023)
par: Yan, Weixiang, et autres
Publié: (2023)
MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings
par: Gurioli, Andrea, et autres
Publié: (2025)
par: Gurioli, Andrea, et autres
Publié: (2025)
Documents similaires
-
MIEB: Massive Image Embedding Benchmark
par: Xiao, Chenghao, et autres
Publié: (2025) -
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
par: Enevoldsen, Kenneth, et autres
Publié: (2024) -
YABLoCo: Yet Another Benchmark for Long Context Code Generation
par: Valeev, Aidar, et autres
Publié: (2025) -
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
par: Han, Hojae, et autres
Publié: (2025) -
Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks
par: Cao, Jialun, et autres
Publié: (2025)