Multimodal Large Language Models (MLLMs) have become a leading approach to visual reasoning, yet the field still lacks a quantitative synthesis that rigorously identifies the drivers of their performance. In this paper, we evaluate 21 models across 7 standardised benchmarks, yielding 102 model–benchmark observations, using a two-level mixed-effects design in which benchmark scores are nested within models. The multilevel model reveals substantial between-model variance (intra-class correlation coefficient, ICC = 0.801) and a significant positive association between model scale and standardised performance (β = 0.73, p = 0.016). Training strategy also matters: relative to instruction tuning, the pretraining + alignment approach is associated with substantially lower performance (β = −1.88, p < 0.001), while early descriptive trends suggest that the Reinforcement Learning from Human Feedback (RLHF) group may be associated with positive effects, although only two models fall within this group, which prevents any definitive conclusion. Correlations between benchmarks are generally strong, ranging from ρ = 0.63 to 0.95 among the reliably estimated pairs. As a sensitivity analysis, we also report a DerSimonian–Laird aggregate meta-analysis, which yields a near-zero pooled effect (d = − 0.046) and extreme heterogeneity (I² = 96.3%). Robustness checks — Hedges’ g validation, leave-one-out sensitivity, bootstrap confidence intervals, SE-floor sensitivity, and Pareto analysis — support the primary conclusions and indicate that training methodology and model scale jointly shape robust multimodal visual reasoning performance.
No takes yet. Share an insight, caveat, or question.
Sharma et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: