Key points are not available for this paper at this time.
Objective: To develop and evaluate an edge-hosted Large Language Model (LLM)-assisted system for automated Neonatal Intensive Care Unit (NICU) discharge summary generation using an evidence-grounded, field-level evaluation framework. Methods: This implementation and evaluation study was conducted in a Level III NICU in India. Longitudinal patient records were constructed from integrated bedside physiologic data (ARCHITECT) and a structured electronic medical record (EMR) platform Although an embedded audio–video module was present, it was not used in this study. Automated discharge summaries were generated by MORPHEUS, an edge-hosted orchestration pipeline running on NVIDIA Jetson AGX Orin hardware with JetPack 6.2. Local orchestration, preprocessing, and workflow execution were performed on the edge device, while language generation inference was performed using the OpenAI gpt-4o-mini API. Documentation quality was assessed with an LLM-based evaluator guided by a clinician-defined rubric comprising 72 fields organized across 14 section contexts and scored on five dimensions: clinical accuracy, completeness, actionability, coherence, and non-hallucination. Paired, field-level comparisons were performed against clinician-authored summaries. Of 549 NICU admissions screened between 1 October 2024 and 3 November 2025, 401 met the inclusion criteria for evaluation. Prompt refinement was performed iteratively using omission-derived feedback without model weight updates. Results: Across 401 evaluated admissions, MORPHEUS-generated summaries demonstrated higher rubric-based scores and lower omission burden than clinician-authored summaries within the structured evaluation framework used in this study, with mean scores of 0.93 versus 0.75 for accuracy, 0.91 versus 0.67 for completeness, 0.93 versus 0.72 for actionability, 0.94 versus 0.74 for coherence, and 0.95 versus 0.78 for non-hallucination, with the largest absolute advantage observed for completeness. Error taxonomy analysis demonstrated fewer omissions, unsupported assertions, and contradictions in AI-generated summaries than in clinician-authored summaries. Iterative prompt refinement was associated with directional improvement across quality dimensions and reduced omission burden, with omission rate per patient decreasing from 2.484 to 1.807 in the later iteration. Conclusions: An edge-hosted LLM-assisted pipeline can generate NICU discharge summaries that meet or exceed clinician-authored documentation quality under a reproducible, clinician-grounded evaluation framework. These findings support the feasibility of deploying edge-orchestrated generative AI systems for high-stakes neonatal clinical documentation using a clinician-grounded field-level evaluation framework.
Singh et al. (Mon,) studied this question.