The Massive Legal Embedding Benchmark (MLEB)
Fuente:
arXiv
Guardado en:
| Autores principales: | Butler, Umar, Butler, Abdur-Rahman, Malec, Adrian Lucas |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Legal RAG Bench: an end-to-end benchmark for legal RAG
por: Butler, Abdur-Rahman, et al.
Publicado: (2026)
por: Butler, Abdur-Rahman, et al.
Publicado: (2026)
MMTEB: Massive Multilingual Text Embedding Benchmark
por: Enevoldsen, Kenneth, et al.
Publicado: (2025)
por: Enevoldsen, Kenneth, et al.
Publicado: (2025)
LimGen: Probing the LLMs for Generating Suggestive Limitations of Research Papers
por: Faizullah, Abdur Rahman Bin Md, et al.
Publicado: (2024)
por: Faizullah, Abdur Rahman Bin Md, et al.
Publicado: (2024)
ALARB: An Arabic Legal Argument Reasoning Benchmark
por: Shairah, Harethah Abu, et al.
Publicado: (2025)
por: Shairah, Harethah Abu, et al.
Publicado: (2025)
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System
por: Su, Weihang, et al.
Publicado: (2025)
por: Su, Weihang, et al.
Publicado: (2025)
MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering
por: Bahaj, Adil, et al.
Publicado: (2025)
por: Bahaj, Adil, et al.
Publicado: (2025)
Interpretable Text Embeddings and Text Similarity Explanation: A Survey
por: Opitz, Juri, et al.
Publicado: (2025)
por: Opitz, Juri, et al.
Publicado: (2025)
Domain-Partitioned Hybrid RAG for Legal Reasoning: Toward Modular and Explainable Legal AI for India
por: Goel, Rakshita, et al.
Publicado: (2025)
por: Goel, Rakshita, et al.
Publicado: (2025)
Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study
por: Han, Jiuzhou, et al.
Publicado: (2024)
por: Han, Jiuzhou, et al.
Publicado: (2024)
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
por: Merrick, Luke, et al.
Publicado: (2024)
por: Merrick, Luke, et al.
Publicado: (2024)
Tabular Embedding Model (TEM): Finetuning Embedding Models For Tabular RAG Applications
por: Khanna, Sujit, et al.
Publicado: (2024)
por: Khanna, Sujit, et al.
Publicado: (2024)
An Ontology-Driven Graph RAG for Legal Norms: A Structural, Temporal, and Deterministic Approach
por: de Martim, Hudson
Publicado: (2025)
por: de Martim, Hudson
Publicado: (2025)
DALDALL: Data Augmentation for Lexical and Semantic Diverse in Legal Domain by leveraging LLM-Persona
por: Choi, Janghyeok, et al.
Publicado: (2026)
por: Choi, Janghyeok, et al.
Publicado: (2026)
Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization
por: Su, Weihang, et al.
Publicado: (2026)
por: Su, Weihang, et al.
Publicado: (2026)
Deterministic Legal Agents: A Canonical Primitive API for Auditable Reasoning over Temporal Knowledge Graphs
por: de Martim, Hudson
Publicado: (2025)
por: de Martim, Hudson
Publicado: (2025)
Predicting Oscar-Nominated Screenplays with Sentence Embeddings
por: Gross, Francis
Publicado: (2025)
por: Gross, Francis
Publicado: (2025)
Quantifying Positional Biases in Text Embedding Models
por: Lee, Reagan J., et al.
Publicado: (2024)
por: Lee, Reagan J., et al.
Publicado: (2024)
Training Sparse Mixture Of Experts Text Embedding Models
por: Nussbaum, Zach, et al.
Publicado: (2025)
por: Nussbaum, Zach, et al.
Publicado: (2025)
Spectral Tempering for Embedding Compression in Dense Passage Retrieval
por: Li, Yongkang, et al.
Publicado: (2026)
por: Li, Yongkang, et al.
Publicado: (2026)
C-Pack: Packed Resources For General Chinese Embeddings
por: Xiao, Shitao, et al.
Publicado: (2023)
por: Xiao, Shitao, et al.
Publicado: (2023)
Length-Induced Embedding Collapse in PLM-based Models
por: Zhou, Yuqi, et al.
Publicado: (2024)
por: Zhou, Yuqi, et al.
Publicado: (2024)
Empowering Meta-Analysis: Leveraging Large Language Models for Scientific Synthesis
por: Ahad, Jawad Ibn, et al.
Publicado: (2024)
por: Ahad, Jawad Ibn, et al.
Publicado: (2024)
Bagging-Based Model Merging for Robust General Text Embeddings
por: Zhang, Hengran, et al.
Publicado: (2026)
por: Zhang, Hengran, et al.
Publicado: (2026)
Little Giants: Synthesizing High-Quality Embedding Data at Scale
por: Chen, Haonan, et al.
Publicado: (2024)
por: Chen, Haonan, et al.
Publicado: (2024)
Enhancing Multilingual Embeddings via Multi-Way Parallel Text Alignment
por: Fazili, Barah, et al.
Publicado: (2026)
por: Fazili, Barah, et al.
Publicado: (2026)
Benchmarking Retrieval-Augmented Generation for Chemistry
por: Zhong, Xianrui, et al.
Publicado: (2025)
por: Zhong, Xianrui, et al.
Publicado: (2025)
NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddings
por: Shachar, Or, et al.
Publicado: (2025)
por: Shachar, Or, et al.
Publicado: (2025)
An Open-Source Dual-Loss Embedding Model for Semantic Retrieval in Higher Education
por: Sajja, Ramteja, et al.
Publicado: (2025)
por: Sajja, Ramteja, et al.
Publicado: (2025)
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
por: Wang, Guangzhi, et al.
Publicado: (2025)
por: Wang, Guangzhi, et al.
Publicado: (2025)
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
por: Sun, Yiqun, et al.
Publicado: (2025)
por: Sun, Yiqun, et al.
Publicado: (2025)
When Text Embedding Meets Large Language Model: A Comprehensive Survey
por: Nie, Zhijie, et al.
Publicado: (2024)
por: Nie, Zhijie, et al.
Publicado: (2024)
Towards Personalized Deep Research: Benchmarks and Evaluations
por: Liang, Yuan, et al.
Publicado: (2025)
por: Liang, Yuan, et al.
Publicado: (2025)
Benchmarking Prompt Sensitivity in Large Language Models
por: Razavi, Amirhossein, et al.
Publicado: (2025)
por: Razavi, Amirhossein, et al.
Publicado: (2025)
Attribution in Scientific Literature: New Benchmark and Methods
por: Saxena, Yash, et al.
Publicado: (2024)
por: Saxena, Yash, et al.
Publicado: (2024)
LegalSeg: Unlocking the Structure of Indian Legal Judgments Through Rhetorical Role Classification
por: Nigam, Shubham Kumar, et al.
Publicado: (2025)
por: Nigam, Shubham Kumar, et al.
Publicado: (2025)
FinMTEB: Finance Massive Text Embedding Benchmark
por: Tang, Yixuan, et al.
Publicado: (2025)
por: Tang, Yixuan, et al.
Publicado: (2025)
ConceptFormer: Towards Efficient Use of Knowledge-Graph Embeddings in Large Language Models
por: Barmettler, Joel, et al.
Publicado: (2025)
por: Barmettler, Joel, et al.
Publicado: (2025)
No Free Lunch in Active Learning: LLM Embedding Quality Dictates Query Strategy Success
por: Rauch, Lukas, et al.
Publicado: (2025)
por: Rauch, Lukas, et al.
Publicado: (2025)
LEMUR: A Corpus for Robust Fine-Tuning of Multilingual Law Embedding Models for Retrieval
por: Ahmadi, Narges Baba, et al.
Publicado: (2026)
por: Ahmadi, Narges Baba, et al.
Publicado: (2026)
One Model Is Enough: Native Retrieval Embeddings from LLM Agent Hidden States
por: Jiang, Bo
Publicado: (2026)
por: Jiang, Bo
Publicado: (2026)
Ejemplares similares
-
Legal RAG Bench: an end-to-end benchmark for legal RAG
por: Butler, Abdur-Rahman, et al.
Publicado: (2026) -
MMTEB: Massive Multilingual Text Embedding Benchmark
por: Enevoldsen, Kenneth, et al.
Publicado: (2025) -
LimGen: Probing the LLMs for Generating Suggestive Limitations of Research Papers
por: Faizullah, Abdur Rahman Bin Md, et al.
Publicado: (2024) -
ALARB: An Arabic Legal Argument Reasoning Benchmark
por: Shairah, Harethah Abu, et al.
Publicado: (2025) -
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System
por: Su, Weihang, et al.
Publicado: (2025)