CodeCSE: A Simple Multilingual Model for Code and Comment Sentence Embeddings

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Varkey, Anthony, Jiang, Siyuan, Huang, Weijing
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914862708293632
author Varkey, Anthony
Jiang, Siyuan
Huang, Weijing
author_facet Varkey, Anthony
Jiang, Siyuan
Huang, Weijing
contents Pretrained language models for code token embeddings are used in code search, code clone detection, and other code-related tasks. Similarly, code function embeddings are useful in such tasks. However, there are no out-of-box models for function embeddings in the current literature. So, this paper proposes CodeCSE, a contrastive learning model that learns embeddings for functions and their descriptions in one space. We evaluated CodeCSE using code search. CodeCSE's multi-lingual zero-shot approach is as efficient as the models finetuned from GraphCodeBERT for specific languages. CodeCSE is open source at https://github.com/emu-se/codecse and the pretrained model is available at the HuggingFace public hub: https://huggingface.co/sjiang1/codecse
format Preprint
id arxiv_https___arxiv_org_abs_2407_06360
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CodeCSE: A Simple Multilingual Model for Code and Comment Sentence Embeddings
Varkey, Anthony
Jiang, Siyuan
Huang, Weijing
Software Engineering
Pretrained language models for code token embeddings are used in code search, code clone detection, and other code-related tasks. Similarly, code function embeddings are useful in such tasks. However, there are no out-of-box models for function embeddings in the current literature. So, this paper proposes CodeCSE, a contrastive learning model that learns embeddings for functions and their descriptions in one space. We evaluated CodeCSE using code search. CodeCSE's multi-lingual zero-shot approach is as efficient as the models finetuned from GraphCodeBERT for specific languages. CodeCSE is open source at https://github.com/emu-se/codecse and the pretrained model is available at the HuggingFace public hub: https://huggingface.co/sjiang1/codecse
title CodeCSE: A Simple Multilingual Model for Code and Comment Sentence Embeddings
topic Software Engineering
url https://arxiv.org/abs/2407.06360