Background: Despite the rapid expansion and volatility of Seattle’s housing market, localized forecasting models for identifying high return-on-investment (ROI) residential properties remain underdeveloped. Prior studies, such as Zhang (2024), have focused on basic regression models using structured data. This study expands upon existing research by integrating ensemble learning techniques and multimodal data sources to enhance ROI prediction accuracy in the Seattle real estate market. Methods: A dataset of over 4,600 residential property transactions in King County, Washington, was retrieved from publicly available data. Structured features included interior space, bedrooms, lot size, ZIP code, school district ratings, and crime rates, while unstructured features were extracted from the property description using transformer-based NLP embeddings. Random Forest, XGBoost, and an ensemble of StackingAveragedModels were trained and compared. Feature selection and optimization were guided by spatial–temporal lag features, SHAP explainability techniques, and Bayesian hyperparameter tuning. Results: StackingAveragedModels yielded the highest results (R² = 0.78, RMSE = USD 88,000, RMSLE = 0.232), outperforming both XGBoost (R² = 0.77) and Random Forest (R² = 0.74). Highest ROI predictors included interior square feet, age of property, schools and transit access, and some linguistic features in descriptions. Conclusions: Integrating engineered property attributes with natural language embeddings through ensemble machine learning enhances ROI forecasting in the urban real estate market significantly. The findings have actionable insights for investors, real estate agents, and urban planners seeking data-informed decisions in high-growth regions like Seattle.
Abhishek Kasturi (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: