We present NPC Nano, a 501M-parameter decoder-only language model pretrained from random initialization on 8.93B tokens using a single NVIDIA A40 GPU. We document the pretraining recipe, a label-shift bug encountered during training and the pre-launch sanity gate that prevents its recurrence, an identity layer methodology with empirically recalibrated capability gates, and a four-experiment characterization of the post-training capacity bottleneck at 0.5B with sub-2% baseline accuracy on GSM8K. None of four post-training interventions (verifiable-reward GRPO, shaped-reward GRPO, self-mined DPO, math-heavy SFT) produced measurable improvement over the SFT baseline, and we argue this convergent failure reflects a capacity bottleneck specific to this scale-baseline regime rather than a property of any individual algorithm. Base and SFT checkpoints are released under Apache 2.0 at huggingface.co/ramankrishna10.
Rama Krishna Bachu (Sat,) studied this question.