Controlled pilot study examines the trade-off between semantic coverage and validity in LLM-generated verification challenges, indicating implications for software testing.
Key Points
The study aims to explore the relationship between semantic coverage and validity in verification challenges generated by large language models.
Conducted a controlled pilot study using LLM to generate verification challenges from software requirements.
Varied access to public examples to measure challenge validity and semantic coverage.
Analyzed results across two distinct software bug families.
Under strong exploration conditions, challenges achieved 30% validity but covered nine useful semantic signatures.
The baseline condition achieved 75% validity with coverage of two useful signatures.
A structured condition reached 100% validity but collapsed to a single semantic signature.