Analyzes political discourse using AI to enhance understanding of multimodal communication.
Today, foundation models simulate humans’ skills in translation, literature review, fact checking, fake-news detection, novel and poetry production. However, generative AI can also be applied to discourse analysis. This study instructed the Gemini 2.5 model to analyze multimodal political discourse. We selected some fragments from the Trump–Zelensky debate held at the White House on 28 February 2025 and annotated each sentence, gesture, intonation, gaze, and facial expression in terms of LEP (Logos, Ethos, Pathos) analysis to assess when speakers, in words or body communication, rely on rational argumentation, stress their own merits or the opponents’ demerits, or express and try to induce emotions in the audience. Through detailed prompts, we asked the Gemini 2.5 model to run the LEP analysis on the same fragments. Then, considering the human’s and model’s annotations in parallel, we proposed a metric to compare their respective analyses and measure discrepancies, finally tuning an optimized prompt for the model’s best performance, which in some cases outperformed the human’s analysis: an interesting application, since the LEP analysis highlights deep aspects of multimodal discourse but is highly time-consuming, while its automatic version allows us to interpret large chunks of speech in a fast but reliable way.
No takes yet. Share an insight, caveat, or question.
Poggi et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: