CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Ye, Meng, Rui, Joty, Shafiq, Savarese, Silvio, Xiong, Caiming, Zhou, Yingbo, Yavuz, Semih |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization
by: Reddy, Revanth Gangi, et al.
Published: (2025)
by: Reddy, Revanth Gangi, et al.
Published: (2025)
SweRank: Software Issue Localization with Code Ranking
by: Reddy, Revanth Gangi, et al.
Published: (2025)
by: Reddy, Revanth Gangi, et al.
Published: (2025)
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
by: Bansal, Srijan, et al.
Published: (2026)
by: Bansal, Srijan, et al.
Published: (2026)
Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement
by: Hayashi, Hiroaki, et al.
Published: (2025)
by: Hayashi, Hiroaki, et al.
Published: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
Modeling Uncertainty and Using Post-fusion as Fallback Improves Retrieval Augmented Generation with LLMs
by: Liu, Ye, et al.
Published: (2023)
by: Liu, Ye, et al.
Published: (2023)
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
by: Niu, Tong, et al.
Published: (2024)
by: Niu, Tong, et al.
Published: (2024)
Automated Classification of Human Code Review Comments with Large Language Models
by: Çağlar, Semih, et al.
Published: (2026)
by: Çağlar, Semih, et al.
Published: (2026)
When More Retrieval Hurts: Retrieval-Augmented Code Review Generation
by: Meng, Qianru, et al.
Published: (2025)
by: Meng, Qianru, et al.
Published: (2025)
Code Review Automation Via Multi-task Federated LLM -- An Empirical Study
by: Kumar, Jahnavi, et al.
Published: (2024)
by: Kumar, Jahnavi, et al.
Published: (2024)
Generating Maximal Configurations and Their Variants Using Code Metrics
by: Yavuz, Tuba, et al.
Published: (2024)
by: Yavuz, Tuba, et al.
Published: (2024)
Contextual Code Retrieval for Commit Message Generation: A Preliminary Study
by: Xiong, Bo, et al.
Published: (2025)
by: Xiong, Bo, et al.
Published: (2025)
PseudoBridge: Pseudo Code as the Bridge for Better Semantic and Logic Alignment in Code Retrieval
by: Li, Yixuan, et al.
Published: (2025)
by: Li, Yixuan, et al.
Published: (2025)
Preference-Guided Refactored Tuning for Retrieval Augmented Code Generation
by: Gao, Xinyu, et al.
Published: (2024)
by: Gao, Xinyu, et al.
Published: (2024)
CodeT5-RNN: Reinforcing Contextual Embeddings for Enhanced Code Comprehension
by: Rahman, Md Mostafizer, et al.
Published: (2026)
by: Rahman, Md Mostafizer, et al.
Published: (2026)
Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Investigating Factuality in Long-Form Text Generation: The Roles of Self-Known and Self-Unknown
by: Tu, Lifu, et al.
Published: (2024)
by: Tu, Lifu, et al.
Published: (2024)
Prompt-based Code Completion via Multi-Retrieval Augmented Generation
by: Tan, Hanzhuo, et al.
Published: (2024)
by: Tan, Hanzhuo, et al.
Published: (2024)
Optimizing Code Embeddings and ML Classifiers for Python Source Code Vulnerability Detection
by: Farasat, Talaya, et al.
Published: (2025)
by: Farasat, Talaya, et al.
Published: (2025)
CodeCSE: A Simple Multilingual Model for Code and Comment Sentence Embeddings
by: Varkey, Anthony, et al.
Published: (2024)
by: Varkey, Anthony, et al.
Published: (2024)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
by: Wang, Sizhe, et al.
Published: (2025)
by: Wang, Sizhe, et al.
Published: (2025)
CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
by: Xu, Jiacheng, et al.
Published: (2025)
by: Xu, Jiacheng, et al.
Published: (2025)
Improving Retrieval-Augmented Code Comment Generation by Retrieving for Generation
by: Lu, Hanzhen, et al.
Published: (2024)
by: Lu, Hanzhen, et al.
Published: (2024)
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
M2CVD: Enhancing Vulnerability Semantic through Multi-Model Collaboration for Code Vulnerability Detection
by: Wang, Ziliang, et al.
Published: (2024)
by: Wang, Ziliang, et al.
Published: (2024)
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale
by: Phan, Huy Nhat, et al.
Published: (2024)
by: Phan, Huy Nhat, et al.
Published: (2024)
Retrieval-Augmented Code Review Comment Generation
by: Hong, Hyunsun, et al.
Published: (2025)
by: Hong, Hyunsun, et al.
Published: (2025)
SCPatcher: Automated Smart Contract Code Repair via Retrieval-Augmented Generation and Knowledge Graph
by: Li, Xiaoqi, et al.
Published: (2026)
by: Li, Xiaoqi, et al.
Published: (2026)
Instructive Code Retriever: Learn from Large Language Model's Feedback for Code Intelligence Tasks
by: Lu, Jiawei, et al.
Published: (2024)
by: Lu, Jiawei, et al.
Published: (2024)
iCodeReviewer: Improving Secure Code Review with Mixture of Prompts
by: Peng, Yun, et al.
Published: (2025)
by: Peng, Yun, et al.
Published: (2025)
Multi-Agent Code-Orchestrated Generation for Reliable Infrastructure-as-Code
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning
by: Tang, Lingxiao, et al.
Published: (2025)
by: Tang, Lingxiao, et al.
Published: (2025)
SERSEM: Selective Entropy-Weighted Scoring for Membership Inference in Code Language Models
by: Dikici, Kıvanç Kuzey, et al.
Published: (2026)
by: Dikici, Kıvanç Kuzey, et al.
Published: (2026)
TypeScript Repository Indexing for Code Agent Retrieval
by: Pu, Junsong, et al.
Published: (2026)
by: Pu, Junsong, et al.
Published: (2026)
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
by: Manh, Dung Nguyen, et al.
Published: (2024)
by: Manh, Dung Nguyen, et al.
Published: (2024)
What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond
by: Gu, Wenchao, et al.
Published: (2025)
by: Gu, Wenchao, et al.
Published: (2025)
Exploring Multi-Lingual Bias of Large Code Models in Code Generation
by: Wang, Chaozheng, et al.
Published: (2024)
by: Wang, Chaozheng, et al.
Published: (2024)
Readability-Robust Code Summarization via Meta Curriculum Learning
by: Zeng, Wenhao, et al.
Published: (2026)
by: Zeng, Wenhao, et al.
Published: (2026)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
by: Wang, Zora Zhiruo, et al.
Published: (2024)
by: Wang, Zora Zhiruo, et al.
Published: (2024)
Similar Items
-
SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization
by: Reddy, Revanth Gangi, et al.
Published: (2025) -
SweRank: Software Issue Localization with Code Ranking
by: Reddy, Revanth Gangi, et al.
Published: (2025) -
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness
by: Le, Hung, et al.
Published: (2024) -
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
by: Bansal, Srijan, et al.
Published: (2026) -
Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement
by: Hayashi, Hiroaki, et al.
Published: (2025)