Key points are not available for this paper at this time.
The large petrochemical melt pump is essential for synthesizing key chemical products in the national economy, with oil monitoring playing a vital role in ensuring the safe operation of its critical gearbox. However, existing deep learning-based intelligent fault detection methods exhibit poor performance under data scarcity conditions, while large language models (LLMs) can identify faults but cannot precisely localize them. To address these problems, we propose a Lightweight Multimodal Vision-Language Model (LMVLM). Firstly, the integrated oil monitoring systems were used to collect online and offline multimodal data (including viscosity, temperature, and wear particle images, etc.) for data preprocessing. Secondly, a lightweight image decoder was designed for pixel-level fault localization using intermediate patch-level features. Additionally, a prompt learner was employed to ensure semantic consistency between the LLM and decoder outputs. Thirdly, the Poisson editing method was utilized to clone objects between images by solving Poisson partial differential equations, with question-answer content being used for prompt tuning. Finally, a multi-objective optimization framework was designed to jointly enhance the LMVLM's performance in semantic alignment, class-imbalance mitigation, and pixel-wise segmentation accuracy. The method was validated using multimodal monitoring data collected from April 2021 to July 2024 on the large petrochemical melt pump gearbox in China. Experimental results demonstrate that LMVLM surpasses conventional deep learning models (CNN, RNN, ResNet-18) in fault detection accuracy and outperforms large language models (GPT-4, GPT-4 Turbo). The proposed research provides an effective solution for intelligent fault detection and localization of critical equipment in the petrochemical industry.
Yang et al. (Mon,) studied this question.