PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 13, 20260 citations

Metrics for Artificial Intelligence in Medicine: A Reference Resource.

View Full Paper
RGRicardo A. GonzalesMTMarcelo Straus TakahashiTRTara Retson

Key Points

  • The aim is to create a standardized framework for evaluating AI metrics in clinical settings.
  • Developed a comprehensive taxonomy of 207 AI performance metrics.
  • Included definitions, citations, synonyms, and mathematical formulae for each metric.
  • Created a structured representation to support reasoning over metric classes.
  • The taxonomy supports evaluation for various data types including structured data and medical images.
  • Logical connections were made to 18 AI model performance criteria, enhancing model comparison.
  • The resource aids in bias detection and assists in selecting appropriate evaluation methods.

Abstract

The effective integration of artificial intelligence (AI) systems into clinical medicine depends on comprehensive and transparent performance evaluation; however, the lack of standardized and widely accepted metrics poses challenges for reproducibility and model adoption. A comprehensive, machine-interpretable framework is presented to formalize the nomenclature and descriptions of 207 graphical, matrix, and scalar metrics used to measure AI model performance. The metrics taxonomy, developed as part of the Radiology Ontology of AI Datasets, Models and Projects (ROADMAP), provides a logically structured representation that captures the semantics of AI evaluation metrics, supports reasoning over metric classes, and enables automated completeness checks for AI model reporting. For each metric, the taxonomy incorporates a definition and citations to authoritative reference sources; where applicable, the taxonomy also includes synonyms, abbreviations, alternate language forms, mathematical formulae, and numerical bounds. The taxonomy supports evaluation of models operating on structured data, medical images, audio signals, and/or unstructured text. Logical axioms link each metric to one or more of 18 AI model performance criteria, including classification, calibration, image segmentation, and text analysis. By harmonizing terminology and enabling structured queries, ROADMAP's taxonomy of AI performance metrics facilitates model comparison, bias detection, and selection of appropriate evaluation methods across diverse datasets and clinical tasks. © RSNA, 2026 See also accompanying Special Report on ROADMAP ontology.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gonzales et al. (2026) studied this question.

synapsesocial.com/papers/69b3ac7002a1e69014cce1fbhttps://doi.org/10.1148/ryai.260070
Ask AI
Helpful
Bookmark
Share
View Full Paper