PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 11, 2026Scientific Reports5 citationsOpen Access

The development and evaluation of agricultural question-answering systems based on large language models

View Full Paper
AEAyşe EldemHEHüseyin Eldem

Key Points

  • To assess the effectiveness of large language models in agricultural question answering.
  • Developed multiple-choice questions across three topics and difficulty levels.
  • Generated answers using GPT-4o and Gemini-2.0-flash LLMs.
  • Applied various prompt strategies like Zero-Shot, CoT, and APE for optimization.
  • Conducted statistical analyses including paired t-tests and ANOVA.
  • LLMs successfully answered questions but showed variability based on model and prompt strategy.
  • Optimized prompts improved consistency in answers.
  • Statistical analysis revealed significant effects based on prompt methods and question difficulty.

Abstract

Large language models (LLMs) show superior performance in different fields. However, the applications of these models, which have shown superior performance in many studies, are still limited and incomplete in agriculture. In this study, a comprehensive evaluation has been made the use of LLMs in the field of agriculture. A set of multiple-choice questions was developed, covering three topics (General, Horticulture, Crop Production) and three difficulty levels (Easy, Medium, Difficult). For each question, answers were generated using GPT-4o and Gemini-2.0-flash LLMs. Zero-Shot, Chain-of-Thought (CoT), Self-Consistency, and Tree-of-Thought (ToT) techniques were preferred as prompt strategies. An automatically optimized prompting pipeline was also applied using Automatic Prompt Engineering (APE) to improve reasoning ability in agricultural question answering. Furthermore, the effect of the prompt methods used on both the accuracy and consistency of the answers was examined. In this study applied in the field of agriculture, the general success of question-answering systems (QAs) was evaluated and the effect of optimized prompts on system success was examined. The findings revealed that LLMs were generally successful, but their results varied significantly depending on the preferred LLM and prompt strategy. The results obtained were analyzed in detail statistically using bootstrap confidence intervals, paired t-tests, ANOVA, and effect size measures (Cohen’s h and d) within the scope of the LLM model, prompt method, difficulty levels, and category-based format. This study introduces one of the first domain-specific question answering systems powered by LLMs, tailored for agricultural experts such as engineers and technicians, and presents an innovative approach by creating an infrastructure for smart digital applications in the field of agriculture.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Eldem et al. (2026) studied this question.

synapsesocial.com/papers/698be001058ab1890a13ba27https://doi.org/10.1038/s41598-026-35003-9
Ask AI
Helpful
Bookmark
Share
View Full Paper