Predicting Future Earnings Changes Using Machine Learning and Detailed Financial Data

Published date01 May 2022
AuthorXI CHEN,YANG HA (TONY) CHO,YIWEI DOU,BARUCH LEV
Date01 May 2022
DOIhttp://doi.org/10.1111/1475-679X.12429
DOI: 10.1111/1475-679X.12429
Journal of Accounting Research
Vol. 60 No. 2 May 2022
Printed in U.S.A.
Predicting Future Earnings Changes
Using Machine Learning and
Detailed Financial Data
XI CHEN,YANG HA (TONY) CHO,YIWEI DOU,
AND BARUCH LEV
Received 30 November 2020; accepted 24 January 2022
ABSTRACT
We use machine learning methods and high-dimensional detailed financial
data to predict the direction of one-year-ahead earnings changes. Our mod-
els show significant out-of-sample predictive power: the area under the re-
ceiver operating characteristics curve ranges from 67.52% to 68.66%, signif-
icantly higher than the 50% of a random guess. The annual size-adjusted
returns to hedge portfolios formed based on the prediction of our models
range from 5.02% to 9.74%. Our models outperform two conventional mod-
els that use logistic regressions and small sets of accounting variables, and
Department of Technology, Operations, and Statistics, Stern School of Business, New York
University; Department of Accounting, Stern School of Business, New York University
February 2022
Accepted by Christian Leuz. An earlier version of this paper circulated under the title “Fun-
damental Analysis of Detailed Financial Data: A Machine Learning Approach.” We benefited
from the comments of an anonymous reviewer, an anonymous associate editor, Aleksander
Aleszczyk, Karthik Balakrishnan, Jeremy Bertomeu, Oliver Binz, Elizabeth Blankespoor, Mark
Bradshaw, John Core, Robert Holthausen, Amy Hutton, Charles M.C. Lee, E. Jin Lee, Becky
Lester, Miao Liu, Joshua Livnat, Michael Minnis, Miguel Minutti-Meza, Joseph Piotroski, K.
Ramesh, Joshua Ronen, Christine Tan, Daniel Taylor, Siew Hong Teoh, Chenqi Zhu, partici-
pants at the 2021 Journal of Accounting Research Conference and the 2021 Transatlantic Doctoral
Conference, and seminar participants at Boston College, Florida International University,New
YorkUniversity, Rice University, UC Irvine, and Stanford University. An online appendix to this
paper can be downloaded at http://research.chicagobooth.edu/arc/journal-of-accounting-
research/online-supplements
467
© 2022 The Chookaszian Accounting Research Center at the University of Chicago Booth School of
Business
468 x. chen, y. h. (t.) cho, y. dou, and b. lev
professional analysts’ forecasts. Analyses suggest that the outperformance rel-
ative to the conventional models stems from both nonlinear predictor inter-
actions missed by regressions and the use of more detailed financial data by
machine learning.
JEL codes: C53, G12, G17, M41
Keywords: direction of earnings changes; prediction; detailed financial
data; XBRL; machine learning
1. Introduction
Developing corporate earnings prediction models is of significant impor-
tance to accounting researchers and investment practitioners. However,
future earnings are difficult to forecast as they are related to numerous
aspects of a company in a complex manner with little guidance from the-
oretical literature (Lev and Gu [2016], Monahan [2018]). Previous studies
often use a small set of financial predictors and regression models. The for-
mer is unlikely to capture the high dimensional aspects relevant to future
earnings; the latter cannot approximate the complex relations. To push
the frontier of earnings prediction, we apply machine learning methods
to a large set of detailed financial data to predict the direction of one-
year-ahead earnings changes. We seek to investigate (1) the out-of-sample
performance of our models and (2) performance differences between our
models and conventional models as well as analysts’ forecasts.
We examine the direction of earnings changes for several reasons. First,
it is difficult to predict the level of future earnings and the amount of earn-
ings changes (future earnings minus known current earnings), as extant
studies find that earnings forecasts based on firm characteristics are not
substantially more accurate than forecasts obtained from the random-walk
model (Gerakos and Gramacy [2013], Li and Mohanram [2014]). Second,
Freeman et al. [1982, p. 643] argue that the variability in earnings changes
is too large to be compared to the variability in expected earnings changes
conditional on explanatory variables. They propose to reduce the variability
in earnings changes by transforming the amount to the direction of earn-
ings changes, predicting which is more achievable. Third, forecasting the
sign of earnings changes is economically meaningful and actionable as ex-
tensive research constructs portfolios based on the direction of earnings
changes (Ou and Penman [1989], Wahlen and Wieland [2011]).
We use two widely accepted machine learning methods based on decision
trees: random forests and stochastic gradient boosting, which have recently
achieved remarkable success in real-world applications (Zhou [2012], Mul-
lainathan and Spiess [2017], Liu [2021]). Compared with regressions, these
methods have three advantages. First, they can accommodate a far more ex-
pansive list of predictors to utilize more nuanced information in detailed fi-
nancial data. For example, we can estimate machine learning models when
the number of predictors is even greater than the number of observations,
predicting future earnings using machine learning 469
whereas traditional regressions break down for such a scenario. Second, the
machine learning algorithms cast a wide net in their specification search to
allow complex associations between high-dimensional predictors and the
predicted variable. Third, these algorithms are specialized for prediction
tasks, rather than explanation tasks. They offer high out-of-sample predic-
tive performance by using the “regularization” (e.g., using a number of
decision trees in random forests) to mitigate overfitting.
To obtain detailed financial data in a machine-readable format, we use
financial reports filed in eXtensible Business Reporting Language (XBRL).
XBRL is an extensible markup language composed of a standard list of tags
(“taxonomy”) to describe business and financial information. Since 2012,
all U.S. public companies must have XBRL tags on quantitative amounts
in financial statements and footnotes of their 10-K reports. Commercial
data aggregators have very limited coverage of these XBRL-tagged detailed
financial data, particularly for footnote disclosures.
Our sample is composed of over 8,000 XBRL filings from 2012 to 2018.
These filings contain more than 4,000 distinct financial items in standard
tags common throughout our sample period. We take all the items for the
current and lagged years, divide them by total assets, and compute the
annual percentage changes, which yield over 12,000 explanatory variables
(i.e., 4,000 ×3 for current values, lagged values, and percentage changes).
For each year in the test period, 2015–2018, we use the second and third
preceding years as the machine learning training period to estimate mod-
els, and the preceding year as the validation period to select the model
that yields the best out-of-sample performance. The chosen model is then
applied to the year in the test period to produce the summary measure
Pr, which characterizes the probability of an increase in the next year’s
earnings.
To evaluate model performance, we use the area under the receiver op-
erating characteristics (ROC) curve (AUC) and 12-month size-adjusted ex-
cess returns to the hedge portfolios formed three months after the fiscal-
year end based on Pr in the test period. While AUC is commonly used
in classification problems, the excess returns offer an economic meaning
for the prediction gains.1We find significant out-of-sample predictability of
our models using machine learning and detailed financial data, concern-
ing the direction of the next year’s earnings changes. The AUC in the test
period ranges from 67.52% to 68.66%, significantly higher than the 50% of
a random guess. The annual size-adjusted returns to the hedge portfolios
1Holthausen and Larcker [1992] use logistic regressions and the same set of financial vari-
ables as Ou and Penman [1989] to directly predict the sign of future stock returns. Recent
studies also apply machine learning to small sets of variables to directly predict future stock re-
turns (Chinco et al. [2019], Rasekhschaffe and Jones [2019], Livnat and Singh [2021]). How
to leverage machine learning methods to analyze a large set of detailed financial data in XBRL
documents for direct return predictions presents an opportunity for future research.

Get this document and AI-powered insights with a free trial of vLex and Vincent AI

Get Started for Free

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex