Machine learning is increasingly used in materials discovery, but its practical application is often limited by the time required to construct structured experimental datasets and by the lack of reliable uncertainty estimates. We introduce an integrated workflow that combines document-grounded large language model (LLM) literature mining, human verification, and probabilistic modeling to enable uncertainty-aware prediction of the coefficient of thermal expansion (CTE) in complex oxides. Experimentally reported CTE values, compositions, and measurement temperature ranges are extracted from full-text articles using a document-grounded large language model and subsequently normalized and verified. Composition-derived descriptors and the reported temperature bounds are used as inputs to a probabilistic regression model that predicts both the expected CTE and a composition-dependent uncertainty, enabling prediction intervals for new compositions. On held-out tests, the model achieves competitive mean accuracy relative to deterministic baselines while producing uncertainty estimates that increase systematically for sparsely represented or chemically distinct compositions, enabling risk-aware screening and prioritization. The workflow supports comparative screening of compositions with targeted CTE behavior and helps guide experimental selection for detailed thermophysical characterization. This study illustrates how LLM-assisted literature curation can be combined with uncertainty-aware machine learning to construct property prediction workflows for materials systems where experimental data are sparse and primarily available in the literature.
Marshall et al. (Mon,) studied this question.