As AI and ML applications become more common in public organisations, technical processes are needed to audit model behaviour, document limitations, and mitigate measurable disparities before deployment. This study presents a methodological framework for fairness auditing and bias mitigation in supervised classification, with emphasis on error disparities across groups rather than a complete solution to AI ethics. The framework is illustrated with two experiments based on data from Chile’s Public Criminal Defence Office (DPP), where the task is to predict whether a criminal case has a favourable or unfavourable outcome for the defendant. Each experiment compares a base model with an improved model that uses reweighting, class weighting, threshold optimisation, interpretability analysis, and documentation. Fairness is operationalised through group-level error metrics, especially false negative rate (FNR), false omission rate (FOR), and true positive rate (TPR), using Region as a geographic proxy attribute. The results show that mitigation can reduce selected disparities, but also reveal a clear fairness–performance trade-off. In the drug trafficking experiment, the manuscript-level confusion matrices show a reduction in FNR from 40.54 to 31.63%, with an increase in FPR from 0.13 to 39.25%. Statistical comparisons are computed from the confusion-matrix counts reported in this manuscript and are therefore reported as aggregate two-proportion comparisons rather than paired McNemar tests. The framework should therefore be read as a technical audit and mitigation procedure, not as evidence that the resulting models are suitable for operational legal decision-making.
Valdebenito et al. (Sat,) studied this question.