The development of new technologies, improved by advances in artificial intelligence, has enabled the creation of a new generation of applications in different scenarios. In medical systems, adopting AI-driven solutions has brought new possibilities, but their effective impacts still need further investigation. In this context, a chatbot prototype trained with large language models (LLMs) was developed using data from the Santa Catarina Telemedicine and Telehealth System (STT) Dermatology module. The system adapts Llama 3 8B via supervised Fine-tuning with QLoRA on a proprietary, domain-specific dataset (33 input-output pairs). Although it achieved 100% Fluency and 89.74% Coherence, Factual Correctness remained low (43.59%), highlighting the limitations of training LLMs on small datasets. In addition to G-Eval metrics, we conducted expert human validation, encompassing both quantitative and qualitative aspects. This low factual score indicates that a retrieval-augmented generation (RAG) mechanism is essential for robust information retrieval, which we outline as a primary direction for future work. This approach enabled a more in-depth analysis of a real-world telemedicine environment, highlighting both the practical challenges and the benefits of implementing LLMs in complex systems, such as those used in telemedicine.
No takes yet. Share an insight, caveat, or question.
Pupo et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: