Key points are not available for this paper at this time.
Purpose: Assessing parental cooperation during child protection services (CPSs) interventions raises distinctive analytical challenges due to ambiguous, conflicting information. We compared reasoning language models (RLMs) to retrieval-augmented generation (RAG) in classifying parental cooperation from Swiss CPS reports. Methods: RLMs of three parameter sizes (255B, 32B, 4B) were benchmarked against a RAG-based approach using a consensus dataset of 100 reports independently classified by two expert human reviewers. Results: The largest RLM achieved the highest accuracy (89%), a nine-point improvement over RAG (80%), approaching expert inter-rater reliability (κ = 0.76). Accuracy was higher for mothers (93%) than fathers (85%), mirrored by expert reviewers. Applied to the full corpus, 31% of cases ( n = 3,900) had at least one parent with documented noncooperation. Discussion : RLMs outperformed RAG and approached expert reliability in classifying parental cooperation. Gender differences likely reflect documentation biases rather than model limitations, corroborating the stronger professional focus on mothers.
Stoll et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: