Abstract Rationale Radiologic interpretation is central to pulmonary sarcoidosis monitoring and research, yet expert assessments (i.e. labels) are known to have high inter-rater variation. Automated image labeling models trained on these assessments are a natural pursuit in light of advancements in machine learning but may inherit assessment limitations, raising concerns about generalizability and clinical utility. We aimed to assess the relationship between reliability of expert labels and generalizability of automated image labeling. Methods Using a single cohort from the GRADS study (N = 368), we analyzed previously obtained chest imaging assessments from four readers / teams of readers with sarcoidosis expertise using four assessment tools producing six distinct assessment contexts (5 CT, 1 CXR). This design allowed us to isolate systematic differences in how readers evaluate images as an obstacle to generalizability of predictive models, distinct from typical focus on differences in population. We evaluated inter-rater agreement using Cohen’s kappa and developed predictive models using demographic and quantitative imaging features to explore systematic differences and generalizability across contexts. We then extracted imaging biomarkers from predictive models and compared their unadjusted associations with pulmonary function testing (PFT) z-scores to those of the radiologic assessments. Results Radiologic assessments exhibited low to moderate agreement across assessment contexts, with 50% of comparisons having kappa below 0.44 and 90% below 0.70. We found strong evidence of assessment differences consistent with noise, radiologist-specific severity thresholds, and varying conceptual interpretation of imaging features. Predictive models trained in one context often failed to generalize to others. Agreement across contexts exhibited correlation of 0.66 with cross-context model AUC, demonstrating reader agreement as a strong limiting factor for generalizable automated image label modelling. By comparison, model-derived imaging biomarkers consistently outperformed original radiologic assessments in explaining PFT variability (see Figure), suggesting potentially greater disease relevance of imaging biomarkers as an alternative to image labels. Conclusions Traditional gold-standard approaches to automated image labeling are heavily undermined by assessment unreliability across contexts. While reliability can be boosted for research through consensus processes, unreliability is currently unavoidable across clinical contexts. As such, results discourage development of gold-standard label models built on expert assessments for clinical purposes until radiologic evaluation standards and reliability improve in sarcoidosis. Imaging biomarkers derived from radiologic assessments offer a more promising and potentially generalizable approach for quantifying disease burden in the meantime. A simple biomarker approach outperformed the original assessments in associating with PFT, a key marker of disease progression in sarcoidosis. This abstract is funded by: NIH, Foundation for Sarcoidosis Research
Lippitt et al. (Fri,) studied this question.