Computational evaluation demonstrates improved question-answering accuracy in electrical customer service, indicating that explicit temporal modeling enhances factual reliability.
Power-system customer service requires answering user queries by jointly interpreting chart images, textual instructions, and high-precision temporal measurements. Existing multimodal large language models are effective at visual-language interaction but remain weak at exact temporal reasoning and reliability-aware decision support. We propose PowerMTQA, a dual-stream multimodal temporal question-answering framework that combines a Qwen-based visual-language backbone with a dedicated multichannel time-series branch, prototype reprogramming, explicit cross-attention fusion, and auxiliary anomaly/risk supervision. Experiments on a unified dataset built from four public electricity and power-related time-series sources show that explicit temporal modeling is critical: dual-stream variants reduce the validation generative loss from 24.5611 for a vision-only baseline to 0.0809 in the single-channel set-ting, while multichannel modeling and LoRA further im-prove the best loss to 0.0721 with 0.9788 token accuracy. On answer-level semantic consistency evaluated by an external judge, PowerMTQA reaches 0.82 accuracy, outperforming the visual-only and text-only baselines by 15 and 39 percentage points, respectively. These results indicate that combining chart understanding with structured temporal reasoning substantially improves factual consistency and practical reliability in power-service question answering.
No takes yet. Share an insight, caveat, or question.
Wang et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: