In the fuzzy front end of innovation, firms often lack sufficient citation, market, and performance data, which limits the usefulness of outcome-based approaches to screening early-stage product innovation opportunities. To address this problem, this study develops a text co-occurrence network-based measurement system for assessing early-stage product innovation opportunities in new product development. We first preprocess idea texts through concept extraction and semantic cleaning, and then construct an integrated semantic network by combining market-related texts with ideation data. The Leiden algorithm is applied to detect latent knowledge communities in the network. Building on this structure, we assess early-stage product innovation opportunities along two complementary dimensions: cross-domain knowledge recombination, capturing the extent to which an idea draws on concept communities that are otherwise weakly connected, and network structural perturbation, capturing the degree to which an idea reconfigures existing semantic boundaries and connection patterns. Based on community entropy and modularity change, we construct a composite indicator for the ex ante assessment of early-stage ideas with stronger product innovation potential. Compared with traditional approaches relying on patent citations, market outcomes, or expert judgments, the proposed method enables earlier screening of ideas that deviate from dominant semantic trajectories and may warrant further development attention. The framework is explicitly positioned as an ex ante screening and attention-allocation tool for early-stage product innovation opportunities, not as a deterministic predictor of later market success.
Wang et al. (Wed,) studied this question.