We present VectraYX-Nano, a 41. 95M-parameter decoder-onlylanguage model trained from scratch in Spanish for cybersecurity, with a Latin-American regional focus and native tool invocationvia the Model Context Protocol (MCP). The model is built aroundfour contributions. (i) Corpus. VectraYX-Sec-ES, a 170M-tokenSpanish corpus assembled by an eight-VM distributed pipeline at∼25 USD of cloud compute and partitioned into three curriculumphases: conversational (42M tokens, OpenSubtitles-ES 35 andOASST1 33), cybersecurity (118M tokens, NVD 39, Wikipedia-ES, in-house NVD-derived Spanish CVE mirror, security blogs), and offensive-security tooling (10M tokens, ExploitDB, HackTricks, OWASP). (ii) Architecture. A 42M-parameter Transformerdecoder combining Grouped-Query Attention 2, QK-Norm 14, RMSNorm 58, SwiGLU 49, RoPE 52, and a 𝑧-loss auxiliary 11, paired with a domain-balanced 16, 384-token byte-fallbackBPE 34, 48 trained on a 50/50 conversational/technical mix-ture. (iii) Curriculum with replay. Continual pre-trainingacross the three phases with a replay buffer 30 mitigatescatastrophic forgetting 19, 32 and yields a monotonic lossdescent (9. 80 → 3. 17 → 3. 00 → 2. 16). After SFT (final loss 1. 74) on a curriculum-aware mixture of OASST-ES, Alpaca-ES, CVEQ a LoRA 29 replication on a260M from-scratch mid-tier reaches 0. 445 ± 0. 201. The releasedGGUF 22 artifact is 81 MB in F16 (approximately 20 MB in4-bit quantization), runs at sub-second time-to-first-token oncommodity hardware under llama. cpp 21, and is, to the best ofour knowledge, the first published Spanish-native cybersecurityLLM with end-to-end MCP integration. We release the corpusconstruction recipe, training scripts, configurations, GGUF weights, ∗ The author is employed at Globant. Institutional affiliation approval is pending. and the B1–B5 benchmark suite (B1: 500, B2: 200, B3: 100, B4: 200, B5: 314 prompts) for reproducibility.
Juan Salas Santillana (Tue,) studied this question.