Gu, Kelly & Xiu perform a comparative analysis of machine-learning methods for the canonical problem of empirical asset pricing — measuring asset risk premiums (the cross-section of expected returns). They demonstrate large economic gains to investors from machine-learning forecasts, in some cases doubling the performance of leading regression-based strategies. The best-performing methods are trees and neural networks, whose predictive gains are traced to allowing nonlinear predictor interactions missed by other methods; all methods agree on the same dominant signals — variations on momentum, liquidity, and volatility.
"We demonstrate large economic gains to investors using machine learning forecasts, in some cases doubling the performance of leading regression-based strategies from the literature. We identify the best-performing methods (trees and neural networks) and trace their predictive gains to allowing nonlinear predictor interactions missed by other methods."
The paper that made machine learning a first-class citizen of empirical asset pricing, and did so carefully: a controlled horse race with honest out-of-sample metrics rather than a single flashy model. Its most durable message is why the flexible methods win — nonlinear interactions among characteristics, not exotic new signals, since every method fingers the same momentum/liquidity/volatility core. That pairs it naturally with Feng-Giglio-Xiu (which prunes the factor list with valid inference) as the two complementary answers to high-dimensional return prediction. The honest caveats are the field's perennial ones — weak, time-varying predictability and black-box interpretability — but the regularize-and-validate template it set is now standard, and it seeds the latent-factor/SDF machine-learning work that followed.