F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911189223604224 |
|---|---|
| author | Zhang, Ziyin Liao, Zihan Yu, Hang Di, Peng Wang, Rui |
| author_facet | Zhang, Ziyin Liao, Zihan Yu, Hang Di, Peng Wang, Rui |
| contents | We introduce F2LLM - Foundation to Feature Large Language Models, a suite of state-of-the-art embedding models in three sizes: 0.6B, 1.7B, and 4B. Unlike previous top-ranking embedding models that require massive contrastive pretraining, sophisticated training pipelines, and costly synthetic training data, F2LLM is directly finetuned from foundation models on 6 million query-document-negative tuples curated from open-source, non-synthetic datasets, striking a strong balance between training cost, model size, and embedding performance. On the MTEB English leaderboard, F2LLM-4B ranks 2nd among models with approximately 4B parameters and 7th overall, while F2LLM-1.7B ranks 1st among models in the 1B-2B size range. To facilitate future research in the field, we release the models, training dataset, and code, positioning F2LLM as a strong, reproducible, and budget-friendly baseline for future works. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_02294 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data Zhang, Ziyin Liao, Zihan Yu, Hang Di, Peng Wang, Rui Computation and Language Artificial Intelligence We introduce F2LLM - Foundation to Feature Large Language Models, a suite of state-of-the-art embedding models in three sizes: 0.6B, 1.7B, and 4B. Unlike previous top-ranking embedding models that require massive contrastive pretraining, sophisticated training pipelines, and costly synthetic training data, F2LLM is directly finetuned from foundation models on 6 million query-document-negative tuples curated from open-source, non-synthetic datasets, striking a strong balance between training cost, model size, and embedding performance. On the MTEB English leaderboard, F2LLM-4B ranks 2nd among models with approximately 4B parameters and 7th overall, while F2LLM-1.7B ranks 1st among models in the 1B-2B size range. To facilitate future research in the field, we release the models, training dataset, and code, positioning F2LLM as a strong, reproducible, and budget-friendly baseline for future works. |
| title | F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.02294 |