Towards Better Understanding of Cybercrime: The Role of Fine-Tuned LLMs in Translation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913295825371136 |
|---|---|
| author | Valeros, Veronica Širokova, Anna Catania, Carlos Garcia, Sebastian |
| author_facet | Valeros, Veronica Širokova, Anna Catania, Carlos Garcia, Sebastian |
| contents | Understanding cybercrime communications is paramount for cybersecurity defence. This often involves translating communications into English for processing, interpreting, and generating timely intelligence. The problem is that translation is hard. Human translation is slow, expensive, and scarce. Machine translation is inaccurate and biased. We propose using fine-tuned Large Language Models (LLM) to generate translations that can accurately capture the nuances of cybercrime language. We apply our technique to public chats from the NoName057(16) Russian-speaking hacktivist group. Our results show that our fine-tuned LLM model is better, faster, more accurate, and able to capture nuances of the language. Our method shows it is possible to achieve high-fidelity translations and significantly reduce costs by a factor ranging from 430 to 23,000 compared to a human translator. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_01940 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Towards Better Understanding of Cybercrime: The Role of Fine-Tuned LLMs in Translation Valeros, Veronica Širokova, Anna Catania, Carlos Garcia, Sebastian Computation and Language Understanding cybercrime communications is paramount for cybersecurity defence. This often involves translating communications into English for processing, interpreting, and generating timely intelligence. The problem is that translation is hard. Human translation is slow, expensive, and scarce. Machine translation is inaccurate and biased. We propose using fine-tuned Large Language Models (LLM) to generate translations that can accurately capture the nuances of cybercrime language. We apply our technique to public chats from the NoName057(16) Russian-speaking hacktivist group. Our results show that our fine-tuned LLM model is better, faster, more accurate, and able to capture nuances of the language. Our method shows it is possible to achieve high-fidelity translations and significantly reduce costs by a factor ranging from 430 to 23,000 compared to a human translator. |
| title | Towards Better Understanding of Cybercrime: The Role of Fine-Tuned LLMs in Translation |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2404.01940 |