KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Xinping, Hu, Xinshuo, Shan, Zifei, Huang, Shouzheng, Zhou, Yao, Zhang, Xin, Sun, Zetian, Liu, Zhenyu, Li, Dongfang, Wei, Xinyuan, Pan, Youcheng, Xiang, Yang, Zhang, Meishan, Wang, Haofen, Yu, Jun, Hu, Baotian, Zhang, Min |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model
di: Hu, Xinshuo, et al.
Pubblicazione: (2025)
di: Hu, Xinshuo, et al.
Pubblicazione: (2025)
On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
di: Zhang, Meishan, et al.
Pubblicazione: (2025)
di: Zhang, Meishan, et al.
Pubblicazione: (2025)
Learning to Extract Rational Evidence via Reinforcement Learning for Retrieval-Augmented Generation
di: Zhao, Xinping, et al.
Pubblicazione: (2025)
di: Zhao, Xinping, et al.
Pubblicazione: (2025)
LMEB: Long-horizon Memory Embedding Benchmark
di: Zhao, Xinping, et al.
Pubblicazione: (2026)
di: Zhao, Xinping, et al.
Pubblicazione: (2026)
In-Context Learning State Vector with Inner and Momentum Optimization
di: Li, Dongfang, et al.
Pubblicazione: (2024)
di: Li, Dongfang, et al.
Pubblicazione: (2024)
CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language Models
di: Li, Dongfang, et al.
Pubblicazione: (2024)
di: Li, Dongfang, et al.
Pubblicazione: (2024)
FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
Take Off the Training Wheels Progressive In-Context Learning for Effective Alignment
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
di: Liu, Zhenyu, et al.
Pubblicazione: (2025)
Improving Attributed Text Generation of Large Language Models via Preference Learning
di: Li, Dongfang, et al.
Pubblicazione: (2024)
di: Li, Dongfang, et al.
Pubblicazione: (2024)
ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution
di: Huang, Shouzheng, et al.
Pubblicazione: (2026)
di: Huang, Shouzheng, et al.
Pubblicazione: (2026)
Is On-Policy Data always the Best Choice for Direct Preference Optimization-based LM Alignment?
di: Sun, Zetian, et al.
Pubblicazione: (2025)
di: Sun, Zetian, et al.
Pubblicazione: (2025)
Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module Operation
di: Hu, Xinshuo, et al.
Pubblicazione: (2023)
di: Hu, Xinshuo, et al.
Pubblicazione: (2023)
Improving Value-based Process Verifier via Low-Cost Variance Reduction
di: Sun, Zetian, et al.
Pubblicazione: (2025)
di: Sun, Zetian, et al.
Pubblicazione: (2025)
Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning
di: Chen, Zhuoen, et al.
Pubblicazione: (2026)
di: Chen, Zhuoen, et al.
Pubblicazione: (2026)
RaSeRec: Retrieval-Augmented Sequential Recommendation
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
Does the Generator Mind its Contexts? An Analysis of Generative Model Faithfulness under Context Transfer
di: Hu, Xinshuo, et al.
Pubblicazione: (2024)
di: Hu, Xinshuo, et al.
Pubblicazione: (2024)
Improving Value-based Process Verifier via Structural Prior Injection
di: Sun, Zetian, et al.
Pubblicazione: (2025)
di: Sun, Zetian, et al.
Pubblicazione: (2025)
KaLM: Knowledge-aligned Autoregressive Language Modeling via Dual-view Knowledge Graph Contrastive Learning
di: Yu, Peng, et al.
Pubblicazione: (2024)
di: Yu, Peng, et al.
Pubblicazione: (2024)
Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training
di: Li, Dongfang, et al.
Pubblicazione: (2023)
di: Li, Dongfang, et al.
Pubblicazione: (2023)
Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards
di: Sun, Zetian, et al.
Pubblicazione: (2025)
di: Sun, Zetian, et al.
Pubblicazione: (2025)
SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
Medico: Towards Hallucination Detection and Correction with Multi-source Evidence Fusion
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
Temporal Knowledge Question Answering via Abstract Reasoning Induction
di: Chen, Ziyang, et al.
Pubblicazione: (2023)
di: Chen, Ziyang, et al.
Pubblicazione: (2023)
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
di: Chen, Huiyao, et al.
Pubblicazione: (2025)
di: Chen, Huiyao, et al.
Pubblicazione: (2025)
From Pre-trained Models to Large Language Models: A Comprehensive Survey of AI-Driven Psychological Computing
di: Chen, Huiyao, et al.
Pubblicazione: (2026)
di: Chen, Huiyao, et al.
Pubblicazione: (2026)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
di: Li, Dongfang, et al.
Pubblicazione: (2026)
di: Li, Dongfang, et al.
Pubblicazione: (2026)
Enhancing Attributed Graph Networks with Alignment and Uniformity Constraints for Session-based Recommendation
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
di: Li, Yunxin, et al.
Pubblicazione: (2025)
di: Li, Yunxin, et al.
Pubblicazione: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
Symplectic Structure-Aware Hamiltonian (Graph) Embeddings
di: Liu, Jiaxu, et al.
Pubblicazione: (2023)
di: Liu, Jiaxu, et al.
Pubblicazione: (2023)
DIVE: Embedding Compression via Self-Limiting Gradient Updates
di: Zhao, Dongfang
Pubblicazione: (2026)
di: Zhao, Dongfang
Pubblicazione: (2026)
Sharp Fractional Sobolev Embeddings on Closed Manifolds
di: Tan, Hao, et al.
Pubblicazione: (2025)
di: Tan, Hao, et al.
Pubblicazione: (2025)
AlphaFuse: Learn ID Embeddings for Sequential Recommendation in Null Space of Language Embeddings
di: Hu, Guoqing, et al.
Pubblicazione: (2025)
di: Hu, Guoqing, et al.
Pubblicazione: (2025)
Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
di: Zhang, Chenyuan, et al.
Pubblicazione: (2026)
di: Zhang, Chenyuan, et al.
Pubblicazione: (2026)
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
di: Lee, Chankyu, et al.
Pubblicazione: (2024)
di: Lee, Chankyu, et al.
Pubblicazione: (2024)
Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection
di: Zeng, Yangchen, et al.
Pubblicazione: (2026)
di: Zeng, Yangchen, et al.
Pubblicazione: (2026)
Approximate Fiber Product: A Preliminary Algebraic-Geometric Perspective on Multimodal Embedding Alignment
di: Zhao, Dongfang
Pubblicazione: (2024)
di: Zhao, Dongfang
Pubblicazione: (2024)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
di: Lin, Gang, et al.
Pubblicazione: (2026)
di: Lin, Gang, et al.
Pubblicazione: (2026)
NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding
di: Liu, Shiyu, et al.
Pubblicazione: (2025)
di: Liu, Shiyu, et al.
Pubblicazione: (2025)
CRoPE: Efficient Parametrization of Rotary Positional Embedding
di: Lou, Beicheng, et al.
Pubblicazione: (2026)
di: Lou, Beicheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model
di: Hu, Xinshuo, et al.
Pubblicazione: (2025) -
On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
di: Zhang, Meishan, et al.
Pubblicazione: (2025) -
Learning to Extract Rational Evidence via Reinforcement Learning for Retrieval-Augmented Generation
di: Zhao, Xinping, et al.
Pubblicazione: (2025) -
LMEB: Long-horizon Memory Embedding Benchmark
di: Zhao, Xinping, et al.
Pubblicazione: (2026) -
In-Context Learning State Vector with Inner and Momentum Optimization
di: Li, Dongfang, et al.
Pubblicazione: (2024)