KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xinping, Hu, Xinshuo, Shan, Zifei, Huang, Shouzheng, Zhou, Yao, Zhang, Xin, Sun, Zetian, Liu, Zhenyu, Li, Dongfang, Wei, Xinyuan, Pan, Youcheng, Xiang, Yang, Zhang, Meishan, Wang, Haofen, Yu, Jun, Hu, Baotian, Zhang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model
by: Hu, Xinshuo, et al.
Published: (2025)
by: Hu, Xinshuo, et al.
Published: (2025)
On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
by: Zhang, Meishan, et al.
Published: (2025)
by: Zhang, Meishan, et al.
Published: (2025)
Learning to Extract Rational Evidence via Reinforcement Learning for Retrieval-Augmented Generation
by: Zhao, Xinping, et al.
Published: (2025)
by: Zhao, Xinping, et al.
Published: (2025)
LMEB: Long-horizon Memory Embedding Benchmark
by: Zhao, Xinping, et al.
Published: (2026)
by: Zhao, Xinping, et al.
Published: (2026)
In-Context Learning State Vector with Inner and Momentum Optimization
by: Li, Dongfang, et al.
Published: (2024)
by: Li, Dongfang, et al.
Published: (2024)
CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language Models
by: Li, Dongfang, et al.
Published: (2024)
by: Li, Dongfang, et al.
Published: (2024)
FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG
by: Zhao, Xinping, et al.
Published: (2024)
by: Zhao, Xinping, et al.
Published: (2024)
Take Off the Training Wheels Progressive In-Context Learning for Effective Alignment
by: Liu, Zhenyu, et al.
Published: (2025)
by: Liu, Zhenyu, et al.
Published: (2025)
Improving Attributed Text Generation of Large Language Models via Preference Learning
by: Li, Dongfang, et al.
Published: (2024)
by: Li, Dongfang, et al.
Published: (2024)
ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution
by: Huang, Shouzheng, et al.
Published: (2026)
by: Huang, Shouzheng, et al.
Published: (2026)
Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module Operation
by: Hu, Xinshuo, et al.
Published: (2023)
by: Hu, Xinshuo, et al.
Published: (2023)
Improving Value-based Process Verifier via Low-Cost Variance Reduction
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning
by: Chen, Zhuoen, et al.
Published: (2026)
by: Chen, Zhuoen, et al.
Published: (2026)
RaSeRec: Retrieval-Augmented Sequential Recommendation
by: Zhao, Xinping, et al.
Published: (2024)
by: Zhao, Xinping, et al.
Published: (2024)
Does the Generator Mind its Contexts? An Analysis of Generative Model Faithfulness under Context Transfer
by: Hu, Xinshuo, et al.
Published: (2024)
by: Hu, Xinshuo, et al.
Published: (2024)
Improving Value-based Process Verifier via Structural Prior Injection
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
KaLM: Knowledge-aligned Autoregressive Language Modeling via Dual-view Knowledge Graph Contrastive Learning
by: Yu, Peng, et al.
Published: (2024)
by: Yu, Peng, et al.
Published: (2024)
Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training
by: Li, Dongfang, et al.
Published: (2023)
by: Li, Dongfang, et al.
Published: (2023)
Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation
by: Zhao, Xinping, et al.
Published: (2024)
by: Zhao, Xinping, et al.
Published: (2024)
Medico: Towards Hallucination Detection and Correction with Multi-source Evidence Fusion
by: Zhao, Xinping, et al.
Published: (2024)
by: Zhao, Xinping, et al.
Published: (2024)
Temporal Knowledge Question Answering via Abstract Reasoning Induction
by: Chen, Ziyang, et al.
Published: (2023)
by: Chen, Ziyang, et al.
Published: (2023)
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
by: Chen, Huiyao, et al.
Published: (2025)
by: Chen, Huiyao, et al.
Published: (2025)
From Pre-trained Models to Large Language Models: A Comprehensive Survey of AI-Driven Psychological Computing
by: Chen, Huiyao, et al.
Published: (2026)
by: Chen, Huiyao, et al.
Published: (2026)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
by: Li, Dongfang, et al.
Published: (2026)
by: Li, Dongfang, et al.
Published: (2026)
Enhancing Attributed Graph Networks with Alignment and Uniformity Constraints for Session-based Recommendation
by: Zhao, Xinping, et al.
Published: (2024)
by: Zhao, Xinping, et al.
Published: (2024)
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
by: Li, Yunxin, et al.
Published: (2025)
by: Li, Yunxin, et al.
Published: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Symplectic Structure-Aware Hamiltonian (Graph) Embeddings
by: Liu, Jiaxu, et al.
Published: (2023)
by: Liu, Jiaxu, et al.
Published: (2023)
DIVE: Embedding Compression via Self-Limiting Gradient Updates
by: Zhao, Dongfang
Published: (2026)
by: Zhao, Dongfang
Published: (2026)
Sharp Fractional Sobolev Embeddings on Closed Manifolds
by: Tan, Hao, et al.
Published: (2025)
by: Tan, Hao, et al.
Published: (2025)
AlphaFuse: Learn ID Embeddings for Sequential Recommendation in Null Space of Language Embeddings
by: Hu, Guoqing, et al.
Published: (2025)
by: Hu, Guoqing, et al.
Published: (2025)
Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
by: Zhang, Chenyuan, et al.
Published: (2026)
by: Zhang, Chenyuan, et al.
Published: (2026)
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
by: Lee, Chankyu, et al.
Published: (2024)
by: Lee, Chankyu, et al.
Published: (2024)
Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection
by: Zeng, Yangchen, et al.
Published: (2026)
by: Zeng, Yangchen, et al.
Published: (2026)
Approximate Fiber Product: A Preliminary Algebraic-Geometric Perspective on Multimodal Embedding Alignment
by: Zhao, Dongfang
Published: (2024)
by: Zhao, Dongfang
Published: (2024)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
by: Lin, Gang, et al.
Published: (2026)
by: Lin, Gang, et al.
Published: (2026)
NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding
by: Liu, Shiyu, et al.
Published: (2025)
by: Liu, Shiyu, et al.
Published: (2025)
CRoPE: Efficient Parametrization of Rotary Positional Embedding
by: Lou, Beicheng, et al.
Published: (2026)
by: Lou, Beicheng, et al.
Published: (2026)
Similar Items
-
KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model
by: Hu, Xinshuo, et al.
Published: (2025) -
On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
by: Zhang, Meishan, et al.
Published: (2025) -
Learning to Extract Rational Evidence via Reinforcement Learning for Retrieval-Augmented Generation
by: Zhao, Xinping, et al.
Published: (2025) -
LMEB: Long-horizon Memory Embedding Benchmark
by: Zhao, Xinping, et al.
Published: (2026) -
In-Context Learning State Vector with Inner and Momentum Optimization
by: Li, Dongfang, et al.
Published: (2024)