In this dissertation, we examine coherence in the context of simplified German-language texts. Coherence is a defining characteristic of a text. It refers to the way the sentences and concepts in a text are connected. Simplification is the act of modifying the content and structure of a text in order to make it easier to understand, whilst still retaining the main content (Alva-Manchego et al., 2020). At a word and sentence level, this could involve replacing words with simpler synonyms or shortening and splitting sentences. At a text level, additional transformations are performed, such as providing background information, or adjusting the structure (cf. Maaß, 2020, p. 89). These adjustments have an effect on the way sentences and concepts are connected. The main research question which this dissertation therefore aims to answer is: which role does coherence play in text complexity for German-language texts? To provide answers to this question, we conduct three empirical analyses: (i) a corpus analysis, (ii) psycholinguistic experiments and (iii) automated approaches to text-level simplification. For our corpus analysis we use a dataset of manually simplified German-language newspaper articles, which we additionally annotate with several layers, including Rhetorical Structure Theory (RST; Mann and Thompson, 1988). Our analysis shows, for example, that the simplified texts contain a less diverse set of RST relations overall. In four self-paced reading experiments we examine if various factors, such as the type of relation or presence of a connective, facilitate comprehension. We find that the presence of a connective results in shorter reading times in concessive relations, compared to causal. We then propose two systems for text-level simplification: feature-based content selection and experiments with a Large Language Model (LLM) for the insertion of information. We evaluate the outputs of these systems, with a focus on coherence. In our manual evaluation, we find that coherence negatively correlates with simplification, suggesting that texts perceived as more ‘simple’ contain more incoherencies. When using automatic metrics to evaluate the coherence, we find that using LLMs to evaluate is a promising approach. However, none of the automatic metrics we tested were able to pick up on seemingly small inconsistencies in the output texts.
Freya Hewett (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: