With the widespread application of the Transformer architecture in various modalities, including vision, the technology of large language models is evolving from a single modality to a multi-modal approach. It is foreseeable that multi-modal large language models will become one of the most powerful tools for humanity in the future.
No takes yet. Share an insight, caveat, or question.
Liang et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: