Experimental evaluation demonstrates superior answer correctness from structured prompting in an agricultural advisory system, indicating that complex retrieval pipelines offer limited benefit.
Agricultural decision-making often depends on timely, evidence-based advice, yet advisory systems for underserved farming communities must also be cost-effective enough to operate continuously. Retrieval-augmented generation (RAG) offers a promising way to provide grounded agricultural recommendations, yet it remains unclear which commonly proposed techniques consistently improve answer quality in real-world deployments. To investigate this question, we developed a bilingual English–Spanish RAG advisory system for Arkansas rice, soybean, and poultry production and used it as a production-oriented testbed under a strict zero-marginal-cost deployment constraint. Through controlled paired experiments evaluated by an independent LLM judge, we compared retrieval strategies, prompt engineering techniques, language models, and evaluation protocols. Several widely used retrieval techniques, including BM25, HyDE, query rewriting, reranking, and token-based chunking, did not improve end-to-end answer correctness over a dense retrieval baseline. In contrast, structured prompting with exemplars consistently improved correctness, while alternative prompt formulations and the freely available language models evaluated offered no measurable advantage over the incumbent 70B model. We also show that conventional single-reference evaluation can substantially underestimate answer correctness in redundant knowledge bases and that independent LLM judging provides a more conservative assessment than self-judging. Together, these results provide practical guidance on which interventions improve answer quality in production-oriented RAG systems and introduce a human-validated multi-reference evaluation protocol for more reliable assessment of answer correctness.
No takes yet. Share an insight, caveat, or question.
Jegede et al. (2026) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: