F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Ziyin, Liao, Zihan, Yu, Hang, Di, Peng, Wang, Rui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
by: Zhang, Ziyin, et al.
Published: (2026)
by: Zhang, Ziyin, et al.
Published: (2026)
F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
by: Zhang, Ziyin, et al.
Published: (2025)
by: Zhang, Ziyin, et al.
Published: (2025)
C2LLM Technical Report: A New Frontier in Code Retrieval via Adaptive Cross-Attention Pooling
by: Qin, Jin, et al.
Published: (2025)
by: Qin, Jin, et al.
Published: (2025)
GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
TingIS: Real-time Risk Event Discovery from Noisy Customer Incidents at Enterprise Scale
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
MELA: Multilingual Evaluation of Linguistic Acceptability
by: Zhang, Ziyin, et al.
Published: (2023)
by: Zhang, Ziyin, et al.
Published: (2023)
Conan-Embedding-v2: Training an LLM from Scratch for Text Embeddings
by: Li, Shiyu, et al.
Published: (2025)
by: Li, Shiyu, et al.
Published: (2025)
Compass-Embedding v4: Robust Contrastive Learning for Multilingual E-commerce Embeddings
by: Ueareeworakul, Pakorn, et al.
Published: (2025)
by: Ueareeworakul, Pakorn, et al.
Published: (2025)
Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code
by: Zhang, Ziyin, et al.
Published: (2023)
by: Zhang, Ziyin, et al.
Published: (2023)
Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs
by: Dou, Longxu, et al.
Published: (2025)
by: Dou, Longxu, et al.
Published: (2025)
SQLfuse: Enhancing Text-to-SQL Performance through Comprehensive LLM Synergy
by: Zhang, Tingkai, et al.
Published: (2024)
by: Zhang, Tingkai, et al.
Published: (2024)
GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation
by: Kim, Yunsu, et al.
Published: (2026)
by: Kim, Yunsu, et al.
Published: (2026)
From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms
by: Jiang, Zhaokun, et al.
Published: (2025)
by: Jiang, Zhaokun, et al.
Published: (2025)
Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages
by: Su, Zeli, et al.
Published: (2025)
by: Su, Zeli, et al.
Published: (2025)
MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings
by: Zhang, Yiqun, et al.
Published: (2026)
by: Zhang, Yiqun, et al.
Published: (2026)
LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
by: He, Ziyuan, et al.
Published: (2025)
by: He, Ziyuan, et al.
Published: (2025)
Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
by: Cheng, Shanbo, et al.
Published: (2025)
by: Cheng, Shanbo, et al.
Published: (2025)
PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
by: Tsai, Yu-Che, et al.
Published: (2026)
by: Tsai, Yu-Che, et al.
Published: (2026)
Measuring Moral LLM Responses in Multilingual Capacities
by: Basu, Kimaya, et al.
Published: (2025)
by: Basu, Kimaya, et al.
Published: (2025)
BayLing 2: A Multilingual Large Language Model with Efficient Language Alignment
by: Zhang, Shaolei, et al.
Published: (2024)
by: Zhang, Shaolei, et al.
Published: (2024)
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
by: Enevoldsen, Kenneth, et al.
Published: (2024)
by: Enevoldsen, Kenneth, et al.
Published: (2024)
SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems
by: Shan, Wenliang, et al.
Published: (2025)
by: Shan, Wenliang, et al.
Published: (2025)
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
by: Zhang, Ziyin, et al.
Published: (2025)
by: Zhang, Ziyin, et al.
Published: (2025)
E^2-LLM: Efficient and Extreme Length Extension of Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment
by: Yang, Wen, et al.
Published: (2025)
by: Yang, Wen, et al.
Published: (2025)
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
by: Tonga, Junior Cedric, et al.
Published: (2025)
by: Tonga, Junior Cedric, et al.
Published: (2025)
Using Optimal Transport as Alignment Objective for fine-tuning Multilingual Contextualized Embeddings
by: Alqahtani, Sawsan, et al.
Published: (2021)
by: Alqahtani, Sawsan, et al.
Published: (2021)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
by: Zhang, Ziyin, et al.
Published: (2024)
by: Zhang, Ziyin, et al.
Published: (2024)
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
by: Sturua, Saba, et al.
Published: (2024)
by: Sturua, Saba, et al.
Published: (2024)
User-LLM: Efficient LLM Contextualization with User Embeddings
by: Ning, Lin, et al.
Published: (2024)
by: Ning, Lin, et al.
Published: (2024)
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
by: Zhang, Hongbin, et al.
Published: (2026)
by: Zhang, Hongbin, et al.
Published: (2026)
FedP$^2$EFT: Federated Learning to Personalize PEFT for Multilingual LLMs
by: Lee, Royson, et al.
Published: (2025)
by: Lee, Royson, et al.
Published: (2025)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
by: Zheng, Weihua, et al.
Published: (2026)
by: Zheng, Weihua, et al.
Published: (2026)
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
by: Günther, Michael, et al.
Published: (2025)
by: Günther, Michael, et al.
Published: (2025)
SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
by: Yu, Zhuohao, et al.
Published: (2025)
by: Yu, Zhuohao, et al.
Published: (2025)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
Incentivizing Inclusive Contributions in Model Sharing Markets
by: Zhang, Enpei, et al.
Published: (2025)
by: Zhang, Enpei, et al.
Published: (2025)
Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation
by: Jiang, Zhaokun, et al.
Published: (2024)
by: Jiang, Zhaokun, et al.
Published: (2024)
ImF: Implicit Fingerprint for Large Language Models
by: Wu, Jiaxuan, et al.
Published: (2025)
by: Wu, Jiaxuan, et al.
Published: (2025)
Similar Items
-
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
by: Zhang, Ziyin, et al.
Published: (2026) -
F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
by: Zhang, Ziyin, et al.
Published: (2025) -
C2LLM Technical Report: A New Frontier in Code Retrieval via Adaptive Cross-Attention Pooling
by: Qin, Jin, et al.
Published: (2025) -
GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding
by: Zhang, Ziyin, et al.
Published: (2024) -
TingIS: Real-time Risk Event Discovery from Noisy Customer Incidents at Enterprise Scale
by: Wang, Jun, et al.
Published: (2026)