The assessment of malingering depression revealed significant performance differences between human evaluators and GPT-3.5 models, highlighting AI's capabilities.
During evaluations, humans accurately identified 75% of cases, whereas GPT-3.5 demonstrated a 65% detection rate in simulated scenarios.
A comparative evaluation was employed, testing both human and AI performance against standardized depression indicators.
Findings suggest that while GPT-3.5 performs well, further validation and improvements are needed for reliable mental health assessments.