Key points are not available for this paper at this time.
Integrating multimodal AI and Intelligent Hyperautomation promises to streamline the understanding of visual artifacts like UML diagrams. We systematically evaluate six leading MLLMs—Gemini, GPT, Claude, Mistral, Grok, and Qwen—on extracting classes, attributes, operations, and relationships from digital UML class diagrams. Our analysis reveals significant performance disparities (F1 scores: 32.1–78.4) and a “class-strong, relationship-weak” pattern. Key failure modes include misinterpreting layout conventions and over-relying on textual proximity. These findings expose architectural limitations in structured visual reasoning, providing actionable insights for future model design, prompting strategies, and hybrid human-AI tooling in model-driven engineering.
Cai et al. (Wed,) studied this question.