PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 18, 20251 citations

Image2Test: Using ChatGPT to Build Manual Tests from Screenshots

View Full Paper
MAManoel ArandaAWAnja WagnerEBEduardo Cristo de Oliveira Barros

Key Points

  • ChatGPT-generated tests showed a 5.56% improvement in completeness compared to human tests, suggesting enhanced thoroughness.
  • 66.7% of ChatGPT tests shared over 50% similarity with human tests, indicating a viable approach for test creation.
  • Human tests were clearer than ChatGPT's by 6.95%, highlighting a trade-off between completeness and clarity.
  • This framework aids in maintaining quality during design changes, potentially saving time and reducing costs.

Abstract

Background:Website layouts often change with new design trends and front-end frameworks. Quality assurance is necessary during these changes, but manual testing takes much time and money. Manual tests are the standard way to maintain quality, but they are slow and expensive. Changes in graphical interfaces can cause errors or break features, which affects quality.Writing manual tests is the most time-consuming part of the process. Aims: This paper presents a tool that uses ChatGPT to create Natural Language Tests from screenshots and operator instructions. The goal is to reduce the time spent on manual test creation and to maintain quality in both the tests and the application. Method:We used two evaluation methods. First, we conducted a survey with 18 software testing professionals and students to compare tests made by ChatGPT and by humans. Second, we used Natural Language Processing techniques to measure the similarity between ChatGPT-generated tests and human-made tests. Results: The qualitative analysis showed that ChatGPT tests exceed human tests in completeness by a difference of 5.56%, achieving 36.11% acceptance rate. Human tests exceeded ChatGPT tests in clarity by 6.95%, reaching 41.67% acceptance rate. The quantitative analysis found that 66.7% of ChatGPT tests shared over 50% similarity with human tests. Conclusions: Our tool can help automate the creation of software tests. The similarity between AI-generated and human-made tests shows that this approach can save time and reduce costs, while keeping test quality at an acceptable level. This framework can help maintain quality during changes in website layouts and application development.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Aranda et al. (2025) studied this question.

synapsesocial.com/papers/68d461bc31b076d99fa60af8https://doi.org/10.5753/sast.2025.14358
Ask AI
Helpful
Bookmark
Share
View Full Paper