Disturbance in language production is a core feature of schizophrenia that has been recognized for more than a century, beginning with Kraepelin's and Bleuler's descriptions of decreases of coherence in spoken language, characterized by derailment and loosening of associations, and a relative poverty of speech. For many decades, the study of language in schizophrenia has remained primarily descriptive, culminating in Andreasen's heuristics of positive (disturbances in coherence) and negative (disturbances in complexity) thought disorder in the 1970s. In the 1980s, Hoffman used a mathematical approach to characterize the misapplication of rules for sentence and discourse formation seen among individuals with schizophrenia, emphasizing semantic relationships between both adjacent and non-adjacent sentences. He developed formal criteria for a "strong hierarchy" of sentences, whose violations would constitute decreases in coherence, that he found to be highly prevalent in schizophrenia1. He subsequently replicated this finding, showing that individuals with schizophrenia have only small or deficient sentence hierarchies, whereas individuals with mania have frequent shifts among large and intact sentence hierarchies. Artificial intelligence was first used to model reduced coherence in speech in the 1990s. Garfield and Rapp2 showed that violations of specific rules in artificial semantic networks could replicate disturbances of spoken language in schizophrenia. Hoffman induced schizophrenia symptoms by reducing connectivity in neural network simulations of parallel, distributed processing systems. Further, he built a computational model of the disorder, finding that deficits in prediction error signaling had the best fit for the pattern of breakdown found in narrative coherence3. In the 21st century, there has been an increasing use of large language models to quantify discourse coherence in schizophrenia. Mota et al4 were creative in the application of speech graphs to transcripts of spoken language, replicating Hoffman's findings with respect to deficits in schizophrenia distinct from those in mania4. In the past decade, the use of automated natural language processing (NLP) to characterize abnormal spoken language in schizophrenia has grown tremendously, extending to include psychosis prediction, such that automated NLP analyses of spoken language have been included in the Accelerating Medicines Partnership in Schizophrenia (AMP SCZ), a large international collaboration that aims to develop multimodal psychosis prediction among at-risk individuals, who notably speak diverse languages5. DISCOURSE in Psychosis is another global initiative launched in 2020 to promote international collaboration in studying language disturbances in psychosis across cultures and languages, with harmonization of methods used to elicit spoken language, so as to create large multilingual datasets for analysis (https://discourseinpsychosis.org). Archived data from these consortia can be used to address key questions in the use of NLP to study language impairment in schizophrenia and its risk states, across languages and cultures. A first question is methodological, i.e., what is the optimal way to elicit language. We use open-ended interviews, to allow for sufficient speech flow to observe decreases in coherence, and to create the ecologically valid context of a dyadic social encounter. This approach also enables the study of other communication modalities in tandem, including acoustic features, pauses, face expression and gesture, with data from both individuals in the dyad being informative. Another key question regards the generalizability of findings across languages and cultures, and how language-specific features may be informative. For example, in a large cross-linguistic study of schizophrenia patients (and controls) – who spoke Danish, German or Chinese – only second-order coherence (i.e., the similarity between phrases separated by another intervening phrase) robustly generalized across languages, while other measures of coherence did not, for unclear reasons that require further study6. In another study of at-risk individuals in Shanghai, both Mandarin-based and English-based NLP methods captured intercorrelated decreases in language-specific coherence (and adjective use), but only Mandarin-based NLP captured greater use of "localizers" (e.g., gongzuo-shang, "during work"; or liangge-ren-zhijian, "between two people") in the at-risk group7. Further, while it is known that the abnormal use of referential noun phrases is common in schizophrenia across languages, this could only be documented in a study conducted in Turkish-speaking patients8. Both AMP SCZ and DISCOURSE in Psychosis offer the opportunity for further cross-linguistic studies, including in other European and Asian languages. These and other datasets also allow to explore the phenomenology of language disturbance in schizophrenia and its risk states beyond coherence and complexity. For example, sentiment analysis can be used to assess the valence and emotional tenor of text. Using a picture description task (positive, negative, neutral) to elicit narrative and speech graph analysis, the connectedness of speech among individuals with first-episode psychosis was found to be directly correlated with the use of positive emotional words. In another study, sentiment analysis was used to identify the emotional tenor of spoken language in open-ended interviews with at-risk individuals, finding greater semantic similarity to "anger" in those who had concurrent suicidal ideation6. In addition to language, the acoustics of spoken language in the schizophrenia spectrum can be assessed, including dysfluencies, timbre/quality, energy/loudness and pause, as well as face expression and gesture in the context of interview6. This yields rich multimodal time series of data that can be used to assess incongruence between different modalities (inappropriate affect) and attunement between conversation partners in language and face expression, indicative not only of psychiatric illness but also of therapeutic alliance. Now we are in the new era of generative large language models (GLMs), which has very significant implications for the study of language (and communication behavior) in schizophrenia. We have used large language models for NLP analyses for fifteen years, and they were silent, but now they can speak. The development of generative artificial intelligence, in particular chatbots, brings forth the notion of language as fundamentally interactive, inter-subjective and cultural. Can GLMs be used not only to measure or quantify features of language, but also to model impairments and then intervene and remediate? In an update to Hoffman's computational approach, GLMs have been used to model the language impairment of schizophrenia. In a recent study8, thought disorder was simulated in narratives using GLMs by increasing the stochasticity of word choice and limiting the model's memory span, with both disturbances decreasing sentence-level coherence. Finally, since self-experience and behavior are constructed through language, language disturbance may represent a constitutive aspect of schizophrenia, and its remediation may be used to treat schizophrenia more broadly. In 1993, Hoffman hypothesized that abnormal discourse planning in schizophrenia leads to both decreased coherence in speech and increased involuntary inner speech. However, he also noted that there were some individuals who heard voices but had coherent speech, albeit relatively simple and rehearsed. He hypothesized that these individuals might compensate for impaired discourse planning by reducing their language output and complexity. He developed a "language therapy" for these "counterexample" patients, which was successful in both improving their discourse planning and reducing their hallucinations9. This is an intriguing but small study for which there have been no efforts at replication. We believe that the use of GLMs may promote new avenues of research on the nature of anomalies of language in schizophrenia, on their meaning as a constitutive aspect of the disorder, and on the possible interventions on these anomalies and consequently more broadly on the disorder.
No takes yet. Share an insight, caveat, or question.
Corcoran et al. (2024) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: