Building on work of Charton, we train small transformer models to calculate the Möbius function (n) and the squarefree indicator function ² (n). The models attain nontrivial predictive power. We apply a mixture of additional models and feature scoring to give a theoretical explanation.
David Lowry-Duda (Thu,) studied this question.