Human and AI collaboration failures and model performance gaps in cardiac surgery: a blinded two-phase evaluation of five large language models | Synapse