PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 11, 2026Journal of Wood Science0 citationsOpen Access

Evaluating fine-tuning and retrieval-augmented generation for domain-specific language modeling in wood science

JBJunsik BangWLWon-Hee LeeBKBonwook Koo

Key Points

  • The central aim is to assess how fine-tuning and retrieval-augmented generation impact the performance of a language model in wood science.
  • Introduced WoodLLaMA, a domain-specific language model fine-tuned on 16,929 wood science articles.
  • Evaluated using two datasets: Journal question-answer set and Wood Handbook QA set.
  • Employed intrinsic metrics (perplexity) and QA-based metrics (cosine similarity, keyword matching, BERTScores).
  • Included qualitative case studies to assess performance.
  • Fine-tuning improved linguistic fluency of the model.
  • Retrieval-augmented generation enhanced semantic alignment.
  • Combining fine-tuning and RAG resulted in the most robust and consistent model performance.
  • Set future directions for utilizing full-text data and integrating human-in-the-loop learning methods.

Abstract

Abstract Recent advances in large language models (LLMs) have produced impressive fluency, yet their application to specialized scientific domains like wood science remains limited. This study introduces WoodLLaMA, a domain-specific LLM fine-tuned on metadata from 16,929 wood science research articles, and examines the effects of fine-tuning and retrieval-augmented generation (RAG) on model performance. Evaluation utilized two datasets not included in the training data: a Journal question–answer (QA) set representing domain-specific expertise and a Wood Handbook QA set reflecting fundamental wood science knowledge. Using intrinsic metrics (perplexity) and QA-based metrics (cosine similarity, keyword matching, and BERTScores), along with qualitative case studies, fine-tuning was found to enhance linguistic fluency while RAG improved semantic alignment. Combining fine-tuning and RAG yielded the most robust and consistent performance. These results demonstrate the complementary value of fine-tuning and RAG for building domain-specific LLMs. The study offers a methodological framework for LLM evaluation and identifies future directions—such as leveraging full-text data, enabling multilingual support, integrating multimodal resources, and incorporating human-in-the-loop learning methods—for enhancing the performance and broadening the applicability of WoodLLaMA across a diverse range of domains.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bang et al. (2026) studied this question.

synapsesocial.com/papers/698c1ca1267fb587c655f293https://doi.org/10.1186/s10086-026-02257-w
Ask AI
Helpful
Bookmark
Share
View Full Paper