This deposit contains the artifacts and documentation for the RPT project, focused on efficient language models with ternary weights (BitNet b1. 58), sparsity, and QAT/STE. The main validated result improves PPL from 25. 13 to 16. 39 (-34. 8%) with 42. 6% sparsity and 100% ternary weights, with functional CPU deployment via GGUF i2ₛ. It also includes notebooks and scripts used to reproduce experiments, as well as logs and validation documentation. Repository and references: - Code and paper: https: //bit. ly/GitHubRPT- Model: https: //bit. ly/rpt-bitnet-2b-pruned- GGUF: https: //bit. ly/rpt-bitnet-2b-pruned-GGUF Note: arXiv submission is in progress (pending endorsement/moderation).
Cesar Favero (2026) studied this question.