Background Computational tools for phylogenetic inference are now routinely applied to data from historical linguistics, especially cognate data. Methods We initially provide an overview of the cognate datasets that are publicly available at present and compare the amount of cognate data with the available masses of molecular data. Then, we outline the drawbacks of the standard binary cognate data representation and introduce an alternative representation that alleviates some of these disadvantages. We also introduce dedicated, parameter-rich evolutionary models for this novel representation. We implement the model and investigate its behavior. In addition, we conduct an orthogonal experiment to investigate whether machine learning-based approaches can be used for cognate data. Results Our experiments show that our newly introduced models can currently not be applied, as they exhibit clear indications for overparameterization due to the small size of the available cognate datasets. We demonstrate that, for the same reason, the applicability of emerging machine learning-based approaches to cognate data is highly limited. Conclusion We conclude that it is necessary to collect more data, investigate potential data sources, and also consider alternative types of data. Historical linguistics will be able to benefit from recent advances in phylogenetics if the amount of available datasets can be substantially increased, both, in terms of number of datasets, and dataset sizes.
Häuser et al. (Mon,) studied this question.