Key points are not available for this paper at this time.
In graphic design, interpreting and transforming abstract content into coherent visual forms is essential. While multimodal large language models (MLLMs) show potential in design generation, they struggle with abstract or imagery-laden input and ignore designers' cognitive processes. Current MLLM-based design tools emphasise resource decomposition and explicit feature matching but lack research on modelling designers' cognitive processes. To address this, this study proposes a design content understanding method that integrates LLM reasoning capabilities with chain-of-thought (CoT) prompting. Drawing on a formative study of how designers select resources, we developed a two-step thinking prompt: the first step elicits a detailed content description, and the second predicts the semantic style values required for a design resource type based on the first-step output. Using a title font selection task in posters as a case study, we compared our method with direct style prediction without CoT prompting, text- or vision-only baselines, and a CLIP-based retrieval baseline. The results demonstrate the effectiveness of the two-step thinking prompt in enhancing design resource selection. These findings provide a reference for prompt design in MLLM-based design tasks and offer insights into optimising design reasoning methods based on large language models, contributing to the development of intelligent design tools.
Sun et al. (Mon,) studied this question.