In recent scientific and regulatory discourse, the terms real-world data (RWD) and real-world evidence (RWE) are often used interchangeably, yet this conflation is conceptually flawed and scientifically consequential. RWD are, by regulatory definition, data routinely collected during the healthcare delivery 1 , 2 . Administrative databases represent one of their most pervasive and scalable forms. Hence, they cover entire populations, allow longitudinal follow-up, and are, in this sense, among the most population-representative real-world sources available. However, as Gale and Hochhaus addressed, the fact that such data are generated in routine practice does not automatically mean that they can support valid inference on the real-world effectiveness of treatment 1 . The translation of data into information, and of information into clinically meaningful knowledge, is not automatic. Failure to recognise this is one of the central methodological challenges in contemporary pharmacoepidemiology and drug utilization studies 3 , 4 , 5 . Health administrative databases are designed primarily for management and governance purposes: they support the reimbursement processes for medications and healthcare services, verify the appropriateness of prescriptions, and monitor the delivery of healthcare services. They were not designed, nor are they structured, to answer questions about therapeutic effectiveness. The abundance of population-level data does not bridge their structural absence of the clinical variables that confer meaning upon what is recorded. Methodologically, this creates a semantic gap : administrative data capture the “when” and the “how much”, but not the “why” or “with what clinical outcome”. This limitation is not a correctable defect; it is inscribed in the very nature of these systems and, for questions of therapeutic effectiveness, cannot be resolved without integration with clinical data, although rigorous drug utilization research remains essential to interpret what administrative data can reliably show 3 , 4 . This limitation can sometimes be mitigated in conditions where therapeutic effectiveness may be reflected in administratively traceable outcomes, such as avoided hospitalisations, or preventable healthcare costs 6 . In chronic myeloid leukaemia (CML), however, this is far less feasible. CML is a rare myeloproliferative neoplasm in which treatment decisions and long-term disease control are guided by serial molecular monitoring rather than by outcomes that routinely generate administrative traces. In patients treated with BCR::ABL1 tyrosine kinase inhibitors (TKIs), clinically decisive response assessments include BCR::ABL1 transcript kinetics, achievement and maintenance of major and deep molecular responses, and eligibility for and outcome of treatment-free remission (TFR) 7 , 8 . These are central to contemporary CML management, yet they do not map reliably onto hospitalisations, diagnostic codes, or other routine administrative markers 7 , 8 . CML therefore represents a particularly informative case in which the gap between RWD and RWE becomes especially clear and empirically testable. This limitation becomes empirically visible when administrative data are interrogated using drug utilization methods rigorously.
Mucherino et al. (Fri,) studied this question.