Purpose This study examines AI–human alignment in thematic analysis from a multi-level perspective, asking whether agreement at the level of individual codes is necessary for convergence in higher-order themes. While prior research often evaluates overall agreement, this study distinguishes between code-level variability and theme-level stability to provide a more nuanced assessment of AI-assisted qualitative analysis. Methods Using qualitative interview data from a doctoral study on e-commerce, the original human thematic analysis (2012, NVivo) was compared with four independent AI-assisted analyses conducted in 2026 using Claude Code. Two prompting strategies were tested: general prompts and structured multi-phase prompts. Alignment was assessed at both code and theme levels using the F1 Score, while inter-session consistency was evaluated through bidirectional mapping across session pairs. Findings At the code level, AI outputs showed considerable variability, with F1 Scores ranging from 55.9% to 87.6% and clear differences between prompting approaches (structured: 83.1% average vs. general: 59.2%). In contrast, theme-level alignment remained consistently high across all sessions (90.9%–100%), with strong inter-session consistency (average 86.1%). These findings indicate that although AI-generated codes may differ across sessions and from human coding, the resulting thematic structures converge reliably. Originality This study introduces a multi-level alignment framework and provides one of the first empirical evaluations of Claude Code as an agentic AI tool for thematic analysis. The 14-year gap between the original human analysis and AI reanalysis offers a distinctive test of AI engagement with historical qualitative data. The study identifies a pattern of hierarchical convergence, where code-level divergence coexists with theme-level stability. Implications The findings suggest that strict code-level agreement may not be necessary for reliable thematic conclusions. AI-assisted analysis can support theme development when structured prompting and human oversight are maintained, offering methodological guidance for integrating AI into qualitative research while preserving analytical rigor.
Rayed AlGhamdi (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: