Tucson Water LCRR Inventory Development, Powered by Artificial IntelligenceAbstractSummary: The Lead and Copper Rule Revision (LCRR) requires that utilities nationwide develop comprehensive service line inventories by October 16, 2024. This presentation focuses on a novel artificial intelligence solution developed in partnership with Microsoft to improve the speed and accuracy of data processing in support of service line inventory development. This solution is being applied with the City of Tucson to efficiently parse data from nearly 100,000 work order comments, extracting valuable information to support development of the City's service line inventory. This presentation will provide an overview of Tucson's approach to LCRR compliance and provide discussion of the enterprise-grade artificial intelligence solution applied for Tucson Water. Abstract: The landscape for regulatory compliance in the water/wastewater industry is undergoing rapid transformation, presenting a host of new challenges for utility managers nationwide. A key example of this evolving regulatory landscape is the Environmental Protection Agency's (EPA) Lead and Copper Rule Revision (LCRR). The objective of LCRR is to reduce the prevalence of lead water service lines in public drinking water systems. Service lines - the pipes that connect water mains to customer homes are typically poorly documented. In many cases, pipe materials for these lines are unknown. To better understand the scope of the lead service line problem, the LCRR requires that utilities nationwide develop structured service line inventories, considering information from all available data sources. The deadline for completion of service line inventories is October 16, 2024. Many utilities lack the required structured data for service connections, but have other data that can be considered, including comment logs from field work orders accumulated over decades of maintenance history. These unstructured text data fields hold a trove of valuable information that can support inventory development and identification of lead service line hotspots. In the case of Tucson Water, the system has 189,000 service lines installed prior to 1988 requiring identification, with nearly 100,000 historical work orders (and associated comment data) that can offer clues to the materials used in these lines. However, the effort to parse this work order data manually via traditional means is prohibitively time consuming and costly. Neither the utility nor consultants have resources available to process this information in time for the LCRR deadline. Recent advances in artificial intelligence - and specifically in large language models (LLMs) have led to new capabilities including strong language understanding, programmability, and fast response times, providing a compelling option to tackle this mountain of data. In collaboration with Microsoft, a custom enterprise-grade LLM solution was developed to parse historical work order data, extract relevant information, and develop LCRR inventories. The solution performs multiple-choice classification and includes algorithmic data-validation to provide quality control. The solution can process approximately 50,000 records per day, with greater speed and precision than manual data entry. The enterprise-grade solution developed in collaboration with Microsoft differs from public large-language models such as Chat GPT in several notable ways. First, the solution is fine-tuned to LCRR data to provide responses in the LCRR-required format. Responses that do not comply with the LCRR required format are automatically rejected. Second, the solution maintains the privacy of the work order data, and data is not stored by the model, nor used for model retraining. Lastly, this solution offers enterprise-grade security and data governance work order data is protected by two-factor authentication and secured by industry leading security standards. Quality assurance and quality control (QA/QC) are of paramount importance when applying artificial intelligence solutions. This project offered the opportunity to provide a mathematical approach to QA/QC. Both precision (repeatability) and accuracy (proximity to correct output) of the program's text classification output were measured mathematically. Model precision was quantified by selecting a statistically significant subset of the data (i.e.: 5,000 classifications) and repeating the model run for this data, creating two independent output datasets and comparing the results. The number of identical responses between the two datasets, divided by the total number of responses, offers a simple quantification of model precision. Model accuracy was quantified by selecting a statistically significant subset of the data (i.e.: 500 classifications), performing manual classification with careful QC, and comparing model output to validated data. Preliminary testing indicates that model precision and accuracy are extremely high in excess of 99% - matching or exceeding the standards required of manual work. This project provides an exciting case study of what is coming next, allowing utility leaders to capture and unlock value from historical data to meet demanding regulatory timelines. For Tucson Water, this AI solution is helping turn nearly 100,000 paragraphs of comment log data into an EPA-compliant LCRR inventory, providing an efficient, effective pathway to regulatory compliance before the October 16, 2024 deadline.This paper was presented at the WEF/AWWA Utility Management Conference, February 13-16, 2024.SpeakerPacker, EricPresentation time14:30:0015:00:00Session time13:30:0015:00:00SessionReal World Applications of Artificial IntelligenceSession number24Session locationOregon Convention Center, Portland, OregonTopicDigital Transformation including AI and ChatGPTTopicDigital Transformation including AI and ChatGPTAuthor(s)Packer, EricAuthor(s)E. Packer1, E. Lansey1Author affiliation(s)HDR 1;SourceProceedings of the Water Environment FederationDocument typeConference PaperPublisherWater Environment FederationPrint publication date Feb 2024DOI10.2175/193864718825159266Volume / Issue Content sourceUtility Management ConferenceWord count10
No takes yet. Share an insight, caveat, or question.
Packer et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: