Effectiveness of large language models in automated evaluation of argumentative essays: finetuning vs. zero-shot prompting | Synapse