CodeCSE: A Simple Multilingual Model for Code and Comment Sentence Embeddings
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866914862708293632 |
|---|---|
| author | Varkey, Anthony Jiang, Siyuan Huang, Weijing |
| author_facet | Varkey, Anthony Jiang, Siyuan Huang, Weijing |
| contents | Pretrained language models for code token embeddings are used in code search, code clone detection, and other code-related tasks. Similarly, code function embeddings are useful in such tasks. However, there are no out-of-box models for function embeddings in the current literature. So, this paper proposes CodeCSE, a contrastive learning model that learns embeddings for functions and their descriptions in one space. We evaluated CodeCSE using code search. CodeCSE's multi-lingual zero-shot approach is as efficient as the models finetuned from GraphCodeBERT for specific languages. CodeCSE is open source at https://github.com/emu-se/codecse and the pretrained model is available at the HuggingFace public hub: https://huggingface.co/sjiang1/codecse |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_06360 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | CodeCSE: A Simple Multilingual Model for Code and Comment Sentence Embeddings Varkey, Anthony Jiang, Siyuan Huang, Weijing Software Engineering Pretrained language models for code token embeddings are used in code search, code clone detection, and other code-related tasks. Similarly, code function embeddings are useful in such tasks. However, there are no out-of-box models for function embeddings in the current literature. So, this paper proposes CodeCSE, a contrastive learning model that learns embeddings for functions and their descriptions in one space. We evaluated CodeCSE using code search. CodeCSE's multi-lingual zero-shot approach is as efficient as the models finetuned from GraphCodeBERT for specific languages. CodeCSE is open source at https://github.com/emu-se/codecse and the pretrained model is available at the HuggingFace public hub: https://huggingface.co/sjiang1/codecse |
| title | CodeCSE: A Simple Multilingual Model for Code and Comment Sentence Embeddings |
| topic | Software Engineering |
| url | https://arxiv.org/abs/2407.06360 |