FinMTEB: Finance Massive Text Embedding Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Yixuan, Yang, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FaMTEB: Massive Text Embedding Benchmark in Persian Language
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025)
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025)
Pooling And Attention: What Are Effective Designs For LLM-Based Embedding Models?
von: Tang, Yixuan, et al.
Veröffentlicht: (2024)
von: Tang, Yixuan, et al.
Veröffentlicht: (2024)
Do We Need Domain-Specific Embedding Models? An Empirical Investigation
von: Tang, Yixuan, et al.
Veröffentlicht: (2024)
von: Tang, Yixuan, et al.
Veröffentlicht: (2024)
MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
von: Ciancone, Mathieu, et al.
Veröffentlicht: (2024)
von: Ciancone, Mathieu, et al.
Veröffentlicht: (2024)
MMTEB: Massive Multilingual Text Embedding Benchmark
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2025)
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2025)
Recent advances in text embedding: A Comprehensive Review of Top-Performing Methods on the MTEB Benchmark
von: Cao, Hongliu
Veröffentlicht: (2024)
von: Cao, Hongliu
Veröffentlicht: (2024)
The Massive Legal Embedding Benchmark (MLEB)
von: Butler, Umar, et al.
Veröffentlicht: (2025)
von: Butler, Umar, et al.
Veröffentlicht: (2025)
JFinTEB: Japanese Financial Text Embedding Benchmark
von: Suzuki, Masahiro, et al.
Veröffentlicht: (2026)
von: Suzuki, Masahiro, et al.
Veröffentlicht: (2026)
Improving Text Embeddings with Large Language Models
von: Wang, Liang, et al.
Veröffentlicht: (2023)
von: Wang, Liang, et al.
Veröffentlicht: (2023)
Rethinking the Privacy of Text Embeddings: A Reproducibility Study of "Text Embeddings Reveal (Almost) As Much As Text"
von: Seputis, Dominykas, et al.
Veröffentlicht: (2025)
von: Seputis, Dominykas, et al.
Veröffentlicht: (2025)
Text Embeddings by Weakly-Supervised Contrastive Pre-training
von: Wang, Liang, et al.
Veröffentlicht: (2022)
von: Wang, Liang, et al.
Veröffentlicht: (2022)
Multilingual E5 Text Embeddings: A Technical Report
von: Wang, Liang, et al.
Veröffentlicht: (2024)
von: Wang, Liang, et al.
Veröffentlicht: (2024)
TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding
von: Qiang, Minjie, et al.
Veröffentlicht: (2026)
von: Qiang, Minjie, et al.
Veröffentlicht: (2026)
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
von: Chen, Jianlyu, et al.
Veröffentlicht: (2025)
von: Chen, Jianlyu, et al.
Veröffentlicht: (2025)
Leveraging Generative Models for Real-Time Query-Driven Text Summarization in Large-Scale Web Search
von: Xiong, Zeyu, et al.
Veröffentlicht: (2025)
von: Xiong, Zeyu, et al.
Veröffentlicht: (2025)
PL-MTEB: Polish Massive Text Embedding Benchmark
von: Poświata, Rafał, et al.
Veröffentlicht: (2024)
von: Poświata, Rafał, et al.
Veröffentlicht: (2024)
ChEmbed: Enhancing Chemical Literature Search Through Domain-Specific Text Embeddings
von: Kasmaee, Ali Shiraee, et al.
Veröffentlicht: (2025)
von: Kasmaee, Ali Shiraee, et al.
Veröffentlicht: (2025)
Enhancing Lexicon-Based Text Embeddings with Large Language Models
von: Lei, Yibin, et al.
Veröffentlicht: (2025)
von: Lei, Yibin, et al.
Veröffentlicht: (2025)
Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
von: Babakhin, Yauhen, et al.
Veröffentlicht: (2025)
von: Babakhin, Yauhen, et al.
Veröffentlicht: (2025)
FinRetrieval: A Benchmark for Financial Data Retrieval by AI Agents
von: Kim, Eric Y., et al.
Veröffentlicht: (2026)
von: Kim, Eric Y., et al.
Veröffentlicht: (2026)
Adapting General-Purpose Embedding Models to Private Datasets Using Keyword-based Retrieval
von: Wei, Yubai, et al.
Veröffentlicht: (2025)
von: Wei, Yubai, et al.
Veröffentlicht: (2025)
Applying Text Embedding Models for Efficient Analysis in Labeled Property Graphs
von: Podstawski, Michal
Veröffentlicht: (2025)
von: Podstawski, Michal
Veröffentlicht: (2025)
Pooling and Semantic Shift: The Fundamental Challenges in Long Text Embedding and Retrieval
von: Gao, Hang, et al.
Veröffentlicht: (2026)
von: Gao, Hang, et al.
Veröffentlicht: (2026)
A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
von: Pham, Loc, et al.
Veröffentlicht: (2025)
von: Pham, Loc, et al.
Veröffentlicht: (2025)
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
von: Hu, Caiyu, et al.
Veröffentlicht: (2025)
von: Hu, Caiyu, et al.
Veröffentlicht: (2025)
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
von: Choi, Chanyeol, et al.
Veröffentlicht: (2025)
JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
von: Pereira, Jayr, et al.
Veröffentlicht: (2026)
von: Pereira, Jayr, et al.
Veröffentlicht: (2026)
FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
Uncovering the Bigger Picture: Comprehensive Event Understanding Via Diverse News Retrieval
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
von: Merrick, Luke, et al.
Veröffentlicht: (2024)
von: Merrick, Luke, et al.
Veröffentlicht: (2024)
Do We Really Need Specialization? Evaluating Generalist Text Embeddings for Zero-Shot Recommendation and Search
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2025)
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2025)
Less is More: Adapting Text Embeddings for Low-Resource Languages with Small Scale Noisy Synthetic Data
von: Navasardyan, Zaruhi, et al.
Veröffentlicht: (2026)
von: Navasardyan, Zaruhi, et al.
Veröffentlicht: (2026)
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
von: Li, Xiangyang, et al.
Veröffentlicht: (2024)
von: Li, Xiangyang, et al.
Veröffentlicht: (2024)
Interpretable Text Embeddings and Text Similarity Explanation: A Survey
von: Opitz, Juri, et al.
Veröffentlicht: (2025)
von: Opitz, Juri, et al.
Veröffentlicht: (2025)
Bagging-Based Model Merging for Robust General Text Embeddings
von: Zhang, Hengran, et al.
Veröffentlicht: (2026)
von: Zhang, Hengran, et al.
Veröffentlicht: (2026)
mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
von: Zhang, Xin, et al.
Veröffentlicht: (2024)
EmbeddingRWKV: State-Centric Retrieval with Reusable States
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
Quantifying Positional Biases in Text Embedding Models
von: Lee, Reagan J., et al.
Veröffentlicht: (2024)
von: Lee, Reagan J., et al.
Veröffentlicht: (2024)
Zero-Shot Contextual Embeddings via Offline Synthetic Corpus Generation
von: Lippmann, Philip, et al.
Veröffentlicht: (2025)
von: Lippmann, Philip, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FaMTEB: Massive Text Embedding Benchmark in Persian Language
von: Zinvandi, Erfan, et al.
Veröffentlicht: (2025) -
Pooling And Attention: What Are Effective Designs For LLM-Based Embedding Models?
von: Tang, Yixuan, et al.
Veröffentlicht: (2024) -
Do We Need Domain-Specific Embedding Models? An Empirical Investigation
von: Tang, Yixuan, et al.
Veröffentlicht: (2024) -
MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
von: Ciancone, Mathieu, et al.
Veröffentlicht: (2024) -
MMTEB: Massive Multilingual Text Embedding Benchmark
von: Enevoldsen, Kenneth, et al.
Veröffentlicht: (2025)