Abstract Background: Patients with high-grade serous ovarian cancer (HGSOC) exhibit highly variable outcomes despite similar treatment protocols, typically involving cytoreductive surgery and platinum-based chemotherapy. Accurate prediction of survival and platinum response immediately after surgery could streamline care pathways and inform early treatment decisions. While computational pathology has focused largely on histology, important prognostic and diagnostic details are often documented in unstructured textual documents, which remain underutilized. This study explores the use of large language models (LLMs) to extract structured representations from multiple textual sources to predict overall survival and platinum response in newly diagnosed HGSOC. Methods: We collected perioperative clinical text, such as pathology reports, radiology reports, history and physical notes, and operative reports from 360 HGSOC patients who underwent cytoreductive surgery followed by platinum-based chemotherapy. Patients with a progression-free interval over 12 months after completion of first-line platinum therapy were classified as responders; those progressing within 12 months as non-responders. Survival was stratified as short (2 years), medium (2–7 years), and long (7 years). Unstructured reports were systematically transformed into structured summaries by addressing expert-curated clinical questions. An LLM extracts concise, structured answers from each document, which are subsequently encoded into embeddings via an embedding model. Machine learning methods then predict survival and platinum response for each individual document type and in a combined multimodal model integrating multiple document types. Model performance was assessed using area under the curve for platinum response, and concordance index for survival using cross-validation. Results: The multimodal model achieved improved metrics on platinum response and survival prediction in comparison to models using traditional tabular clinical data or computational pathology-derived features. Models trained on single document types performed worse highlighting the benefit of integrating a variety of documentation. Feature attribution maps reveal consistent predictive signals from operative complexity and extent of disease. Conclusions: We demonstrate the utility of LLMs to convert fragmented clinical narratives into structured data for prognosis in HGSOC. By enabling accurate upfront prediction of platinum response and survival using only routine clinical documentation, this approach offers a scalable method to personalize treatment strategies. Future work will focus on incorporating additional data modalities and performing external validation to evaluate the model’s generalizability across institutions and languages. Citation Format: Farieda Gaber, Leonhard Donle, Altuna Akalin, Ernst Lengyel. Application of large language models for prognostic prediction in high-grade serous ovarian cancer using unstructured clinical text abstract. In: Proceedings of the AACR Special Conference in Cancer Research: Advances in Ovarian Cancer Research; 2025 Sep 19-21; Denver, CO. Philadelphia (PA): AACR; Cancer Res 2025;85 (18Suppl): Abstract nr A028.
Gaber et al. (Fri,) studied this question.