Key points are not available for this paper at this time.
Recent years have seen rapid progress in multimodal large language models (MLLMs) in the field of chemistry. However, chemical MLLMs capable of handling cross-modal understanding and generation remain underexplored. To fill this gap, we propose ChemMLLM, a unified chemical MLLM for molecule understanding and generation. In this work, we design five types of multimodal tasks across text, molecular simplified molecular input line entry system (SMILES) strings, and images and curate the datasets. We benchmark ChemMLLM against a range of general leading MLLMs, chemical LLMs, and specialized models on these tasks. Experimental results demonstrate the feasibility of a unified multimodal interface. Our work extends the capabilities of chemical MLLMs to the realm of image generation, demonstrating the feasibility of unifying multiple cross-modal chemical tasks within a single mode l and enabling more intuitive, visual human-artificial intelligence (AI) interaction.
Tan et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: