The F-measure, also known as the F1-score, is widely used to assess the performance of classification algorithms. However, some researchers find it lacking in intuitive interpretation, questioning the appropriateness of combining two aspects of performance as conceptually distinct as precision and recall, and also questioning whether the harmonic mean is the best way to combine them. To ease this concern, we describe a simple transformation of the F-measure, which we call F^* F ∗ (F-star), which has an immediate practical interpretation.
No takes yet. Share an insight, caveat, or question.
Hand et al. (2021) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: