Key points are not available for this paper at this time.
Rice crop residue burning (RCRB) in India constitutes a major annual environmental crisis, contributing significantly to regional air pollution, greenhouse gas emissions, and public health deterioration across the Indo-Gangetic Plain. Despite growing policy attention, a systematic, data-driven understanding of the diverse perspectives—agricultural, environmental, economic, and socio-political—expressed across multiple textual sources remains lacking. This study proposes a large language model (LLM)-driven topic modeling pipeline leveraging TopicGPT, an instruction-tuned prompting framework, to extract and evaluate high-level thematic insights from heterogeneous text corpora related to RCRB in India. Our pipeline integrates four sequential stages—topic generation, topic refinement, multi-label topic assignment with grounded evidence, and assignment correction—operated via a locally deployed LLM through the Ollama inference framework. Post-extraction, we evaluate topic quality using ten quantitative metrics encompassing embedding-based coherence, inter-topic diversity, Silhouette score, Davies–Bouldin index, Calinski–Harabasz score, and distribution entropy, among others. Results demonstrate that the proposed pipeline effectively recovers semantically coherent and diverse topic structures from multi-source text data, offering actionable insights for policymakers and researchers addressing RCRB.
Naganawa et al. (Thu,) studied this question.