VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity LLM with Native MCP Tool Integration
Fuente:
Zenodo
Enregistré dans:
| Auteur principal: | |
|---|---|
| Format: | Recurso digital |
| Langue: | anglais |
| Publié: |
Zenodo
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866902089545809920 |
|---|---|
| author | Salas Santillana, Juan |
| author_facet | Salas Santillana, Juan |
| contents | <p>We present VectraYX-Nano, a 41.95M-parameter decoder-only language model trained<br>from scratch in Spanish for cybersecurity, with a Latin-American regional focus<br>and native tool invocation via the Model Context Protocol (MCP). The model is<br>built around four contributions. (i) Corpus: VectraYX-Sec-ES, a 170M-token<br>Spanish corpus assembled by an eight-VM distributed pipeline at ~$25 USD of cloud<br>compute, partitioned into three curriculum phases: conversational (42M tokens,<br>OpenSubtitles-ES and OASST1), cybersecurity (118M tokens, NVD, Wikipedia-ES,<br>in-house Spanish CVE mirror, security blogs), and offensive-security tooling<br>(10M tokens, ExploitDB, HackTricks, OWASP). (ii) Architecture: a 42M-parameter<br>Transformer decoder combining Grouped-Query Attention, QK-Norm, RMSNorm, SwiGLU,<br>RoPE, and a z-loss auxiliary, paired with a domain-balanced 16,384-token<br>byte-fallback BPE tokenizer. (iii) Curriculum with replay: continual pre-training<br>across three phases with a replay buffer mitigates catastrophic forgetting and<br>yields a monotonic loss descent (9.80 -> 3.17 -> 3.00 -> 2.16). After SFT the<br>released model attains a conversational gate of 0.78 +/- 0.05 over N=4 seeds.<br>(iv) Two empirical findings: A controlled bootstrap-corpus ablation exposes a<br>loss-versus-register inversion -- lower-perplexity bootstraps yield worse<br>conversational behavior at the nano scale. A post-hoc LoRA study shows that the<br>tool-selection floor of 0.000 on the mixed SFT corpus is a corpus-density<br>artifact: a tool-dense corpus (2,801 examples, ratio 1:21) raises B4 to<br>0.145 +/- 0.046 on Nano 42M and 0.445 +/- 0.201 on a 260M mid-tier (N=4 seeds).<br>The released GGUF artifact is 81 MB in F16 (~20 MB in 4-bit), runs at<br>sub-second time-to-first-token on commodity hardware, and is, to our knowledge,<br>the first published Spanish-native cybersecurity LLM with end-to-end MCP<br>integration.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_20122226 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity LLM with Native MCP Tool Integration Salas Santillana, Juan Spanish NLP cybersecurity LLM MCP tool use curriculum pre-training LATAM <p>We present VectraYX-Nano, a 41.95M-parameter decoder-only language model trained<br>from scratch in Spanish for cybersecurity, with a Latin-American regional focus<br>and native tool invocation via the Model Context Protocol (MCP). The model is<br>built around four contributions. (i) Corpus: VectraYX-Sec-ES, a 170M-token<br>Spanish corpus assembled by an eight-VM distributed pipeline at ~$25 USD of cloud<br>compute, partitioned into three curriculum phases: conversational (42M tokens,<br>OpenSubtitles-ES and OASST1), cybersecurity (118M tokens, NVD, Wikipedia-ES,<br>in-house Spanish CVE mirror, security blogs), and offensive-security tooling<br>(10M tokens, ExploitDB, HackTricks, OWASP). (ii) Architecture: a 42M-parameter<br>Transformer decoder combining Grouped-Query Attention, QK-Norm, RMSNorm, SwiGLU,<br>RoPE, and a z-loss auxiliary, paired with a domain-balanced 16,384-token<br>byte-fallback BPE tokenizer. (iii) Curriculum with replay: continual pre-training<br>across three phases with a replay buffer mitigates catastrophic forgetting and<br>yields a monotonic loss descent (9.80 -> 3.17 -> 3.00 -> 2.16). After SFT the<br>released model attains a conversational gate of 0.78 +/- 0.05 over N=4 seeds.<br>(iv) Two empirical findings: A controlled bootstrap-corpus ablation exposes a<br>loss-versus-register inversion -- lower-perplexity bootstraps yield worse<br>conversational behavior at the nano scale. A post-hoc LoRA study shows that the<br>tool-selection floor of 0.000 on the mixed SFT corpus is a corpus-density<br>artifact: a tool-dense corpus (2,801 examples, ratio 1:21) raises B4 to<br>0.145 +/- 0.046 on Nano 42M and 0.445 +/- 0.201 on a 260M mid-tier (N=4 seeds).<br>The released GGUF artifact is 81 MB in F16 (~20 MB in 4-bit), runs at<br>sub-second time-to-first-token on commodity hardware, and is, to our knowledge,<br>the first published Spanish-native cybersecurity LLM with end-to-end MCP<br>integration.</p> |
| title | VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity LLM with Native MCP Tool Integration |
| topic | Spanish NLP cybersecurity LLM MCP tool use curriculum pre-training LATAM |
| url | https://doi.org/10.5281/zenodo.20122226 |