Context: Automated compliance testing faces challenges when requirements demand semantic understanding. Traditional tools evaluate at most 50% of compliance criteria, necessitating expensive manual verification. Objectives: This study evaluates Large Language Model (LLM) capabilities for web accessibility conformance testing across 39 WCAG 2 AA success criteria, examining performance, prompt engineering, and proposing an interpretive framework. Methods: We evaluated seven cost-effective LLMs on isolated HTML snippets from 384 W3C ACT cases. Three prompting strategies were tested with five replications, comparing LLMs against traditional tools. Results characterise snippet-level evaluation, which may differ from full-page testing. Results: The best configuration (deepseek-reasoner under provider-default settings) achieved 71.52% accuracy with example-based prompts—an 8.2% cross-model average improvement over basic prompts. Traditional tools achieved 22.9–33.9% real accuracy (treating coverage gaps as errors) and 22.9–44.4% decision accuracy. LLMs classified every test case (100% coverage), reflecting their always-answering design rather than superior capability, whereas tools selectively abstain. LLMs struggled with determining test applicability (45%–67% error rates) and Understandable criteria. Conclusion: Contributions include: (1) an unprecedented 39-criteria empirical evaluation; (2) a dual-metric framework (Real vs. Decision Accuracy) clarifying the structural difference between always-answering LLMs and abstaining tools; (3) evidence that prompt engineering improves performance; and (4) a post-hoc interpretive framework (based on expert judgment, not validated measurement) relating criterion characteristics to LLM amenability to generate future hypotheses. LLMs remain best suited for human-in-the-loop workflows.
López-Gil et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: