PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 6, 2026International Journal of Information and Learning Technology0 citations

Testing large language models as a simulation tool for multi-theory analysis to simulate research outcomes

View Full Paper
CCChi-Kuan ChiaLBLukas BaschungARAhmad Zabidi Abdul Razak

Key Points

  • To investigate the use of large language models as simulation tools for enhancing doctoral research efficiency.
  • Conducted theory simulations using GPT-5, DeepSeek-V3.2, and Mistral Small 3.1.
  • Performed a quantitative evaluation with independent raters calculating scores and deviation indices.
  • Implemented a thematic analysis of the simulated outputs.
  • GPT-5 achieved the highest percentage maximum possible score compared to other models.
  • The average deviation index was below the critical value, confirming inter-rater reliability.
  • Two of the three models produced similar outputs, indicating consistency in results.

Abstract

Purpose This study aims to test large language models as a tool for multi-theory simulation that could potentially augment the doctoral student’s (DS) research efficiency. Design/methodology/approach The study proceeded in three phases, commencing with theory simulations where identical prompts were applied to GPT-5, DeepSeek-V3.2 and Mistral Small 3.1. Phase Two involved a quantitative evaluation by two independent raters using the percentage of maximum possible score and the average deviation index. Finally, Phase Three entailed a thematic analysis of the simulated outputs. Findings Large language models can augment research efficiency by simulating varied theoretical perspectives and assisting the student in assessing the suitability of theory before commencing a real-world study. Quantitative analysis showed that GPT-5 achieved the highest percentage maximum possible score, followed by DeepSeek-V3.2, and finally Mistral Small 3.1. Additionally, the average deviation index was 0.53, which is less than the critical value of 0.67. Therefore, the inter-rater reliability was established. The qualitative analysis revealed that two of the three models generated the same output, while the third generated a slightly different output. Originality/value This research integrates large language model simulation into educational research methodology before conducting empirical studies. This method could potentially augment doctoral research efficiency and present the DS with novel perspectives on their research areas.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chia et al. (2026) studied this question.

synapsesocial.com/papers/69faa2b504f884e66b533542https://doi.org/10.1108/ijilt-09-2025-0270
Ask AI
Helpful
Bookmark
Share
View Full Paper